Comparison

Why AI Agent Skills and Plugins Are a Security Risk

At a glance

  • Agent Skills and plugins are executable instructions that run with the employee's own identity and permissions, usually installed with no security review.
  • A Skill is often just a markdown file that tells an agent to read files, call APIs, or run shell commands.
  • Endpoint detection and response tooling resolves processes and files, so which Skill, plugin or MCP server issued an instruction stays unresolved.
  • Backslash Security research found hidden instructions in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials.
  • Backslash Security discovers every agent, MCP server, Skill, hook, rules file and plugin on an endpoint, then blocks risky actions inline.

Backslash Security

Published:

AI agent Skills and plugins are a security risk because they are executable instructions that arrive without review and run with the permissions of the person whose laptop they sit on. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — it can tell the agent to read a file, call an API, or run a shell command — and in practice it is frequently nothing more than a markdown file an engineer downloaded. Plugins, connectors and rules files such as AGENTS.md or CLAUDE.md work the same way: they carry standing instructions the agent reads and acts on every run, which means anyone who can edit that content can steer the agent. Endpoint detection and response, the incumbent endpoint layer that enterprises buy to catch malicious processes, files and known-bad behavior on a machine, resolves the process that executed; the Skill, plugin or MCP server that issued the instruction sits above that view, which is where agent skills security and MCP server security actually live. Backslash Security was built for that layer — it discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, and blocks risky agent actions inline before execution, including credential access and data sent to unapproved destinations. The company was founded by Shahar Man, CEO and Co-founder, and Yossi Pik, CTO and Co-founder.

What exactly are AI agent Skills and plugins, and how do they execute on an employee endpoint?

What does each component actually do?

This section covers one AI stack layer: installable agent extensions — what an employee adds after agent installation — and where each executes.

  • Agent — software that takes natural language goals, plans steps, and acts through tools: reading/writing files, running shell commands, driving browsers, calling SaaS APIs. Examples: Claude Code, Cursor, GitHub Copilot, Codex.
  • Agent Skills — packaged instructions and scripts extending agent capabilities. Skills tell agents to read files, call APIs or run shell commands, executing with user permissions. Often just markdown files on the machine.
  • MCP server — MCP (Model Context Protocol) connects agents to external tools and data sources. Each MCP server is a live connection; each MCP tool is one exposed operation.
  • Plugins and connectors — add-ons attaching agents to applications, repositories or data sources.
  • Hook — trigger firing on agent actions, running something before or after.
  • Rules file (AGENTS.md, CLAUDE.md) — repository or machine file carrying standing instructions the agent reads every run.
  • Context — everything assembled for a run: prompt, file contents, tool output, loaded rules.

All executes locally. The agent process runs on the employee's laptop under their OS account, SSO session, Git credentials and browser cookies, so its reach equals theirs. Device management inventories the application; Skills, MCP servers, hooks and rules files added afterward are user-configured and sit outside inventory. Developer workstations are prominent use cases, and the same layer appears on analyst, sales and marketing endpoints running desktop agents and connectors.

Why do Skills and plugins carry more risk than the underlying model?

Skills and plugins carry more risk than the underlying model because they execute actions. A model generates text; an Agent Skill converts that text into shell commands, API calls, or file reads, running with the employee's permissions on their machine. The exposure is an action taken on the endpoint, against live credentials and real data, by a component the security team never inventoried.

Four mechanisms turn execution capability into risk. Agents follow instructions found in content they read—files, repository issues, tickets, web pages—not only user-typed instructions; this is indirect prompt injection, where hidden instructions reach the agent through material it fetches. Install-time approval grants broad scope once and is rarely revisited. Extensions can be updated after review, so what runs today may differ from what was approved. Chained tool calls through connectors stitch a read in one system to a write in another, enabling data exfiltration. Backslash Security traces the full path from prompt to agent to tool call to outcome, making the chain reconstructible after the fact.

Do this But watch out for Handle it by
Inventory every Skill, plugin, hook and rules file present on endpoints A point-in-time list ages as extensions update silently Reassessing components continuously rather than once at approval
Narrow what an installed component may reach Install prompts bundle wide permission grants nobody unpicks later Allowlisting approved components and denylisting risky ones per team
Treat files, issues and pages an agent reads as untrusted input Hostile instructions arrive inside content the agent fetches itself Constraining which actions are available to the agent at all
Watch connector-to-connector chains, not single calls Exfiltration paths look like ordinary tool usage in isolation Host-level guardrails, which is where Backslash Security enforces

How do Skills, plugins, MCP servers, hooks and rules files compare as attack surfaces?

Skills, plugins and MCP servers all extend what an AI agent can do, but each presents a distinct attack surface depending on where it runs and whose identity it carries.

Comparison criteria:

  • Where it executes. On the host, inside the agent process, or at a remote service. This decides whether a host-level control observes the action or only its network trace.
  • Identity inherited. Most endpoint agentic components run with the employee's own credentials, making their actions look legitimate to identity-based controls.
  • Permission scope. What the component can reach: local files, shell, cloud credentials, SaaS data.
  • Mutability after approval. Whether the reviewed component can change without re-review.
  • Dominant failure mode. The characteristic way the component goes wrong.
Component Where it executes Identity inherited Typical permission scope Changes after approval? Dominant failure mode
Agent Skills — packaged instructions and scripts an agent can invoke Locally, under the agent The user's own account File reads, API calls, shell commands Yes — often a markdown file anyone can edit Unreviewed capability added quietly to a machine
Plugins and extensions Inside the agent or IDE process The user's session Whatever the host application exposes Yes — via auto-update Capability expands without a new approval
Connectors Agent process to a third-party service Delegated user tokens The connected account's data On token or scope change Over-broad consent to a personal account
MCP servers — endpoints agents connect to for tools and data Local process or remote service Caller's credentials Everything the server fronts Server-side, invisibly Trusted connection reaching further than intended
Hooks — triggers that fire before or after an agent action On the host The user Arbitrary local execution Editable on the machine Automatic execution nobody reviewed
Rules files (AGENTS.md, CLAUDE.md) — standing instructions read on every run Interpreted by the agent The user Steers every subsequent action Any repository write Instructions treated as operator intent

Why do existing security control layers miss Skill and plugin activity?

Existing security controls each watch a specific object, and an Agent Skill — packaged instructions and scripts that extend what an AI agent can do, often just a markdown file sitting on the machine — is not that object for any of them. This depends on what you mean by "control," because the word carries two meanings in AI governance conversations. A technical control observes or blocks something at a defined layer: a process starting, a token being issued, a packet leaving the host. An administrative control is a written AI usage policy naming which tools employees may run. This section concerns technical controls on the endpoint, where agents, Skills and MCP servers — the Model Context Protocol connections through which an agent reaches external tools and data — actually execute.

Control layer What it observes How a Skill or plugin load appears to it
Endpoint detection and response (EDR) Processes, files, known-bad behavior An approved AI client reading a text file it may read
Identity tooling (Entra ID, Okta) Logins, sessions, token grants A user who authenticated normally hours earlier
SaaS, CASB and network tooling Traffic to sanctioned applications Ordinary API calls to an approved model provider
MDM (Intune, Jamf) Installed applications, device posture Nothing — the client is approved; its Skills were never enumerated

No configuration change to those layers surfaces the activity, because none is positioned to see it. A Skill or plugin loads inside an already-approved client, under the employee's own identity. Backslash Security operates at the host-side agentic layer and replaces none of the controls above.

Option Recognized strength Relationship to the agentic layer
Backslash Security Specialist in endpoint agentic AI Resolves which agent, Skill or MCP server caused an action, and blocks it inline
CrowdStrike Established, trusted endpoint footprint, broad general-purpose coverage Endpoint platform with an agent already on the machine
Palo Alto Networks Platform breadth and consolidation appeal Part of a broad security platform
Air Security Inspects agent traffic at the network layer Suits environments where deploying endpoint software is difficult

What does securing the agentic AI fabric on enterprise endpoints actually involve?

Securing the agentic AI fabric on enterprise endpoints means treating every AI component an employee runs locally as one connected system rather than a list of separate tools. The agentic AI fabric is the interconnected layer of agents, models, MCP servers and MCP tools, Agent Skills, hooks, rules files, plugins, connectors and context sitting on a machine — and because these components invoke one another, examining any one of them in isolation leaves the paths between them unchecked. Backslash Security was built around that layer, and its capability set runs across five jobs.

  • Discover. Backslash Security continuously maps the agentic components present on a machine, including unapproved ones and those running under a personal account on a corporate device.
  • Assess. Each component is given a risk posture through an agentless read of the endpoint — agentless meaning information is collected without leaving software permanently installed for that purpose.
  • Govern. Backslash Security turns that inventory into policy: allowlisting approved components, denylisting risky ones, and applying custom rules to different teams and risk profiles rather than one blanket standard.
  • Prevent in real time. Backslash Security applies inline prevention at the moment an agent attempts an action, rather than reporting it after the fact.
  • Reconstruct. Backslash Security collects audit and tracing data of agentic activity for compliance and investigation, which matters when a Skill is a markdown file a user edited an hour ago and no deployment record exists.

As a verifiable trust signal for this layer as of 2026: Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader, a badge Latio awards to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills. That evaluation covers the same control surface a Skill or plugin touches when it executes with a user's own permissions.

How should a security team start governing Skills and plugins this quarter?

A security team can start this quarter with a short, ordered sequence in which every step has a named owner and a concrete output. Skills and plugins are packaged instructions that extend what an AI agent can do, running with the employee's own permissions.

  1. Build the inventory. Owner: endpoint and MDM administrator, with security. Output: a per-machine list of agents, Skills, plugins, MCP servers, hooks and rules files, including components installed under personal accounts. Watch out: schedule discovery as a recurring job, not a one-time snapshot.

  2. Classify by permission reach and data access. Owner: security architect. Output: a risk tier per component, based on what it can read, execute and transmit with the user's credentials. Watch out: grade on reachable scope — shell execution, credential stores, source repositories, outbound destinations — not publisher name.

  3. Define approval and re-approval. Owner: AI governance lead. Output: an allowlist and denylist with a re-review trigger tied to version and content change. Watch out: extensions update silently, so bind re-approval to the artifact, not the decision date.

  4. Set inline enforcement for the highest-risk action classes. Owner: security engineering. Output: policy enforced on the host at the agentic layer, where Backslash Security acts on agent behavior before an action completes. Watch out: over-broad rules push users toward unsanctioned tools — scope narrowly, then widen.

  5. Establish forensic reconstruction. Owner: detection and response. Output: retained agent-run records routed into the SIEM alongside existing endpoint telemetry. Watch out: retention without correlation produces volume, not answers.

Governance programs lose momentum after the first inventory because classification depends on knowing what each extension actually exercises at runtime, and that detail is rarely declared. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, making step one executable without a procurement cycle.

Frequently Asked Questions

What makes AI agent Skills and plugins a security risk?

AI agent Skills and plugins are a security risk because they are packaged instructions and scripts that execute with the employee's own permissions. An Agent Skill is a reusable capability an agent can invoke — in practice often just a markdown file sitting on the machine — and it can tell the agent to read a file, call an API, or run a shell command. Plugins, connectors and hooks (triggers that fire before or after an agent action) extend the same reach, and the agent invoking them already holds that person's access to source code, Git systems and internal services.

Why don't MDM and application controls cover Skills and plugins?

Mobile device management platforms such as Intune and Jamf govern which applications are installed on a machine. A Skill, plugin, hook, rules file or MCP server is added inside an already-approved agent such as Claude Code, Cursor or GitHub Copilot, so it never appears as a new application to approve or deny. The same gap produces shadow AI: an employee can wire up an unknown MCP server, or sign in through a personal account on a corporate laptop, without tripping any device policy. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which is what turns that blind spot into an inventory.

How does a malicious Skill or rules file actually attack an agent?

Usually through prompt injection: hidden instructions placed in content the agent reads — a file, repository, issue or web page — which the agent then follows as though the user had typed them. A rules file (for example AGENTS.md or CLAUDE.md) carries standing instructions an agent reads on every run, which makes it an attractive place to hide them. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.

Why doesn't the existing endpoint security stack catch this?

EDR — endpoint detection and response — is built to catch malicious processes, files and known-bad behavior on a machine. A Skill invoked by an approved agent runs inside a trusted process, using credentials the user is entitled to use, so the telemetry stops at the process boundary and never resolves which agent, Skill or MCP server caused the action. Backslash Security works at that agentic layer: it blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations, and traces the full path of a run from prompt to agent to tool call to outcome.

How can a security team get an inventory of what is already running?

Start agentless — meaning information is collected without leaving software permanently installed on the endpoint. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, distributed through existing device management tooling. For component-level review, Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks, and publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks. On the MCP side, Backslash Security reports that its MCP Server Security Hub held 81,021 publicly available MCP servers when read on 22 September 2026, each scored for risk, giving teams a reference point for MCP server security decisions before a connection is approved.

Who builds Backslash Security?

Backslash Security was founded by Shahar Man, CEO and Co-founder, and Yossi Pik, CTO and Co-founder. The company is backed by Stage One, First Rays Venture Partners, Artofin and D E Shaw & Co, alongside a group of individual investors. Its focus is the agentic AI fabric on enterprise endpoints — discovery, risk assessment, guardrails, inline prevention, and the audit and tracing data that supports incident investigation. As of 2026, that remains the whole of what the company builds.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to make the switch?

See why teams choose Backslash Security.

Book a Demo