At a glance
- Most AI plugin review fails at install time, then never revisits the agents, MCP servers, Skills, hooks and rules files that change afterward.
- An Agent Skill is often just a markdown file, and it executes with the employee's own permissions and access.
- Latio's 2026 AI Security Market Report calls Backslash Security "one of the most complete options available," covering Claude Code, Cursor, MCP servers and Skills.
- Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.
Backslash Security
Published:
Security teams vetting AI plugins tend to go wrong in five specific places: reviewing the marketplace listing rather than the installed artifact; treating a read-only description as a read-only permission; approving the plugin but not the tool chain and context sources it reaches, such as MCP servers, Agent Skills, hooks and rules files; assuming installation is a one-time event that never needs reassessing; and leaving no record of what the plugin actually did. Those five failures matter because of what is actually being installed as of 2026, where the item submitted for approval may just as easily be an MCP server — Model Context Protocol being the protocol agents use to reach external tools and data, with each server a live connection an agent can act through — as a conventional add-on. An Agent Skill is a packaged set of instructions and scripts that extends what an agent can do, frequently a markdown file that tells the agent to read a file, call an API or run a shell command under the user's own permissions. A rules file such as AGENTS.md or CLAUDE.md carries standing instructions the agent reads on every run, and a hook is a trigger that fires before or after an agent action. The stakes are concrete: Backslash Security research discovered that malicious repository configurations can redirect Claude Code API requests to an attacker-controlled server, silently exposing the user's sign-in token and allowing attackers to run AI workloads, access profile information, upload files, and initiate sessions under the victim's identity. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, assesses what is safe to use, and blocks risky agent actions inline before execution.
What counts as an AI plugin, and why is vetting one different from vetting ordinary software?
What counts as an AI plugin on an employee's machine is broader than the word implies, and the distinction matters because each carries instructions and permissions, not just features. This section covers only components that live on the endpoint and run with that employee's identity and access, where traditional software review breaks down.
| Component | What it is | What it actually carries |
|---|---|---|
| Plugin / extension | An add-on that extends an AI client such as Cursor or Claude Desktop | Code plus permissions inherited from the host application |
| Connector | A configured link between an agent and a data source or SaaS system | Standing credentials and a reachable destination |
| MCP server | A service reachable over the Model Context Protocol, the protocol agents use to connect to external tools and data | A connection an agent can act through, with its own tool surface |
| MCP tool | A single callable action exposed by an MCP server | A described capability the agent chooses to invoke on its own |
| Skill | Packaged instructions and scripts that extend what an agent can do — in practice often a markdown file on the machine | Instructions that execute with the user's own permissions |
| Hook | A trigger that fires on an agent action, running something before or after it | Automatic execution outside the prompt the user typed |
| Rules file (AGENTS.md, CLAUDE.md) | A file in a repository or on a machine carrying standing instructions an agent reads on every run | Persistent directives applied to every session |
Conventional application review asks for a vendor, a version, a signature and a known-vulnerability record. Several items above have none of those properties: a Skill or rules file is text an agent reads and obeys, so its risk lives in what it instructs and what it can reach. A review process built around installed binaries will return a clean result for a markdown file that tells an agent to read a credentials directory, because nothing in that file looks like software. Backslash Security treats these components as the unit of assessment, scoring each MCP server and Skill for risk and judging the combinations between them, since a safe Skill and a safe MCP server can still form a problematic pairing.
Which five mistakes most often slip past an AI plugin review?
Five mistakes recur most often when security teams vet AI plugins, Agent Skills and MCP servers. An Agent Skill is packaged instructions and scripts that extend an agent and execute with the user's own permissions; an MCP server is a Model Context Protocol connection an agent acts through.
| Mistake | Do this instead | But watch out for |
|---|---|---|
| Reviewing the marketplace listing rather than the installed artifact | Read the files on disk: manifest, scripts, bundled commands | Publishers ship new versions after approval, so re-read the artifact, never the listing |
| Treating a read-only description as a read-only permission | Map what the component can actually invoke — shell, filesystem, network destinations | Declared scope and granted scope diverge quietly; verify at the permission layer |
| Approving the plugin but not its tool chain and context sources | Enumerate the MCP servers, connectors and rules files it reaches | A rules file such as AGENTS.md or CLAUDE.md carries standing instructions the agent obeys on every run |
| Assuming install is a one-time event | Treat discovery as continuous and re-assess on change | Employees install under their own identities, so point-in-time inventory decays quickly |
| Leaving no record of what the plugin actually did | Capture prompt, tool call and outcome for every run | Without tracing, an incident becomes reconstruction from memory |
What makes context sources so dangerous?
The agent reads them as instructions. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — a file no plugin review would think to open. The component itself was never modified, which is why artifact inspection alone does not close the gap.
How do you prove what a plugin actually did afterward?
Through tracing at the agent layer rather than the process layer. An endpoint log shows that an interpreter ran and a network call left the machine; it does not show which prompt triggered the run, which tool the agent selected, or what came back. Reconstructing an incident requires that chain to be recorded while the run happens, because the evidence does not exist anywhere else afterward.
Why does a one-time approval stop being true after the plugin updates itself?
A one-time approval stops being true the moment the thing you approved changes underneath the record. Approval captures a component as it existed on the day it was reviewed; the agentic layer on an endpoint does not hold still. Through 2026, the common drift paths are: a plugin auto-updates from its distribution channel, an MCP server publishes a new manifest that registers additional tools, a remotely hosted server changes behavior with no local version bump, a rules file such as AGENTS.md or CLAUDE.md is edited after review, or the repository behind an approved component transfers to new owners.
| Do this | But watch out for — and how to cover it |
|---|---|
| Pin and record versions of approved plugins and MCP servers | Remotely hosted servers and prompt-side updates change behavior with no version change; re-verify the live component at run time, not just the manifest on file |
| Re-read manifests whenever they change | New tool registrations can appear inside an already-approved server; diff the registered tool set and its permissions, not the server name |
| Monitor ownership of source repositories and packages | A transfer hands an approved supply chain to an unknown maintainer; trigger reassessment on maintainer or endpoint change, not on a review calendar |
| Treat rules files, hooks and Skills as mutable configuration | A standing instruction edited post-approval alters every subsequent run; track file content changes and enforce at the point of action |
Continuous reassessment needs an inventory that refreshes itself rather than a sign-off filed once. Backslash Security continuously discovers the agentic components on an endpoint, assesses their risk posture, and enforces allowlist and denylist policy continuously, with custom policies for different teams and risk profiles.
Which vetting checks hold up at review time, and which only work at execution time?
Vetting checks hold up at review time for the properties a component declares before it runs: what it is, where it came from, and what permissions it requests. What an agent actually does once a prompt, repository file, or tool response enters the loop is only observable at execution time. The two layers fail in different ways.
The review-time layer covers inventory of agentic components on a machine—agents, MCP servers, Skills, hooks and rules files, meaning files carrying standing instructions an agent reads on every run—plus assessment of manifests and requested permissions, publisher provenance, and allowlist or denylist policy decisions. The execution-time layer evaluates the specific action about to run, blocks it inline before execution, and records the run for later reconstruction.
Four criteria separate the layers:
- Field of view—what the layer can observe, which sets the class of risk it can catch.
- Latency—how long between a change on the endpoint and a control responding, which decides whether a risky action is prevented or merely logged.
- Failure mode—how the control behaves when wrong, since false negatives and false positives carry different costs.
- Evidence produced—what artifact remains afterward for audit, incident response, or compliance.
| Criterion | Review-time layer | Execution-time layer |
|---|---|---|
| Can see | Declared permissions, manifest contents, provenance, component inventory | The concrete action requested, its arguments, destination, and run context |
| Cannot see | Runtime intent, injected instructions reaching the agent mid-run | Components never invoked during an observed session |
| Latency | Point-in-time; drifts as files and configurations change | Immediate, evaluated before the action completes |
| Failure mode | A benign-looking Skill—often just a markdown file—passes and misbehaves later | Over-blocking disrupts legitimate work; narrow policy lets an action through |
| Evidence | Approval record and risk score at decision time | Trace of prompt, agent, tool call and outcome |
Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
How do employee identities and permission sprawl hide plugin risk on endpoints?
When a plugin, agent or connector is installed by an employee rather than provisioned centrally, it runs with that person's identities, tokens and permission grants instead of a scoped service account. The blast radius follows the human: OAuth consents to SaaS apps, cloud CLI credentials, Git configuration, browser sessions. Permission sprawl accrues to the person rather than the tool, so each approval silently widens what every later component can reach.
Three mechanics do most of the hiding. Credentials are shared by proximity—a token written into a config file is readable by any other agent on the same machine. Access scope creeps because an MCP server—Model Context Protocol, the protocol agents use to connect to external tools and data—stays connected long after the task that justified it. And Agent Skills, packaged instructions and scripts often just a markdown file, execute with the user's own permissions, with no separate grant to review.
| Attribute | Possible values | Why it decides risk |
|---|---|---|
| Running identity | Corporate SSO (Okta, Entra ID), personal account, local user | A personal login is shadow AI in its most concrete form: unsanctioned use security cannot see or revoke |
| Credential source | OS keychain, .env file, Git config, long-lived API token | Determines what an agent reaches if it reads the wrong file |
| Access scope | Single repository, whole SaaS tenant, production environment | Sets how far one misjudged approval travels |
| Component type | Agent, MCP server, Skill, hook, rules file, plugin, connector | Each is granted and revoked differently |
Directory and device-management platforms such as Entra ID, Okta, Intune or Jamf enumerate managed applications and sanctioned logins; none sees a component a user installed locally under a personal account. Backslash Security works at the layer where these components actually execute, so a vetting decision can rest on which identity is running which component today. Developer workstations are a prominent instance of this pattern, and the same identity inheritance appears wherever staff run assistants under their own logins.
Frequently Asked Questions
What actually counts as an "AI plugin" when security teams vet AI plugins?
Security teams vetting AI plugins usually picture a browser extension or an IDE add-on, but on a modern endpoint the installable surface is much wider. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint — the components an approval workflow has to cover:
- Agent Skills — packaged instructions and scripts that extend what an agent can do; a Skill can tell an agent to read a file, call an API or run a shell command, and it executes with the user's own permissions. In practice a Skill is often just a markdown file on the machine.
- MCP servers — connections made through the Model Context Protocol, the protocol agents use to reach external tools and data sources. Each server is a path an agent can act through.
- Rules files such as AGENTS.md or CLAUDE.md — files in a repository or on a machine carrying standing instructions an agent reads on every run.
- Hooks — triggers that fire on an agent action, running something before or after it.
Why isn't a one-time approval review enough for AI plugins?
A one-time review certifies a component as it looked on the day it was read, while the agentic layer on an endpoint changes continuously: Skills get added, MCP servers get swapped, rules files get edited, and an approved agent can still become a rogue agent — reaching for credentials, escalating privilege or sending data somewhere nobody approved — after being manipulated or drifting from its task. Backslash Security addresses this with continuous discovery and enforcement instead of a point-in-time check: allowlisting of approved components, denylisting of risky ones, and custom policies tuned to different teams and risk profiles.
How can a file an agent reads become an attack on the endpoint?
Through prompt injection: hidden instructions placed in content an agent reads — a file, repository, issue, web page or document — which the agent then follows as though the user had typed them. The indirect form matters most on the endpoint, because the agent reaches the poisoned content on its own, without anyone pasting anything. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Separately, Backslash Security research discovered that malicious repository configurations can redirect Claude Code API requests to an attacker-controlled server, silently exposing the user's sign-in token — allowing attackers to run AI workloads, access profile information, upload files, and initiate sessions under the victim's identity. Neither case requires malware on the machine, which is why file-level and process-level review passes them.
Can existing endpoint and MDM tooling govern this layer?
Largely no, and that gap is the reason plugin vetting stalls. EDR — endpoint detection and response, the incumbent endpoint layer — is built to catch malicious processes, files and known-bad behavior on a machine; an agent's prompt, its tool calls and the instructions it was given sit above that boundary. MDM governs device configuration and software distribution, so it has no view of which Skills or MCP servers an agent installed inside a sanctioned application. Backslash Security works at the agentic layer itself and blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
How can a team build an inventory without installing anything on endpoints?
By running an agentless assessment — collecting information without leaving software permanently installed on the machine. Backslash Security offers a free AI Endpoint Exposure Assessment, called Scout, that is agentless, read-only, and retains no data; it is distributed through the organization's existing MDM, runs, reports and disappears. That read surfaces shadow AI — the agents, MCP servers and Skills nobody approved — including the concrete case of an employee connecting through a private account on a corporate machine. Agentless applies to the discovery step; continuous enforcement and inline prevention are a separate capability.
Which agents does the coverage span, and what audit evidence comes out of it?
Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity, across all employees and all endpoints, with developer workstations as a prominent use case. Latio's 2026 AI Security Market Report concludes that Backslash Security is "one of the most complete options available," with coverage spanning the full agentic surface from Claude Code and Cursor to the MCP servers and Skills running on top of them. On the evidence side, Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and per Backslash Security it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC — the record a team needs to reconstruct an incident or answer an auditor.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07