Comparison

Mistakes Security Teams Make When Vetting AI Agent Skills

At a glance

  • Most Skill reviews are one-time file reads, yet an Agent Skill is an editable markdown file that can change after approval.
  • A Skill executes with the employee's own permissions, so vetting that stops at "is it malicious?" misses credential access and data egress.
  • Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.
  • Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader for in-depth controls over AI on the endpoint.

Backslash Security

Published:

Security teams vetting AI agent Skills repeat a short list of mistakes: they review a Skill once at approval time and never again, they read it as a document instead of as executable capability, they judge it in isolation from the agent, MCP servers and rules files it operates beside, and they have no inventory telling them which Skills are actually installed on employee machines in the first place. An Agent Skill is packaged instructions and scripts that extend what an AI agent can do — in practice often just a markdown file on a laptop — and it can tell an agent to read a file, call an API or run a shell command, executing with that employee's own permissions and access. That last property is why a clean read-through proves less than reviewers assume: nothing in the file is signed, nothing prevents an edit the day after approval, and the agent that loads it answers to the user's identity, not to a service account the security team governs. As of 2026, Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity, the agents through which Skills typically reach an endpoint. The sections below work through each vetting mistake and what a durable review process does differently.

What exactly is an AI agent Skill, and why does vetting one differ from vetting an app?

An AI agent Skill is, exactly, a packaged capability an agent loads to carry out a task: instructions, tool bindings and context bundled together — in practice often just a markdown file on an employee's machine, sometimes with scripts beside it. Vetting one differs from vetting an application because there is no installer, no signed binary and no process name to allowlist. A Skill can tell an agent to read a file, call an API or run a shell command, and it executes with the user's own permissions.

How does a Skill differ from the components around it?

Component What it is Where its risk shows up
Agent Skill Packaged instructions and scripts the agent loads to do a task The actions it authorizes once invoked
MCP server A connection to external tools or data via the Model Context Protocol The systems it exposes to the agent
MCP tool An individual callable operation offered by an MCP server The specific operation invoked, not the server as a whole
Hook A trigger that runs something before or after an agent action Code executing outside the agent's visible task
Rules file (AGENTS.md / CLAUDE.md) Standing instructions the agent reads on every run Persistent influence over behavior, applied silently
Plugin / connector An extension binding the agent to an application or service Standing access granted under the user's identity

Which Skill attributes belong in the record?

  • Invocation trigger — automatic, model-selected or user-called; determines whether anyone sees it run.
  • Effective permissions — inherited from the signed-in employee, including shell, filesystem and cloud credentials.
  • External reach — the APIs, endpoints and MCP servers it can touch.
  • Mutability — whether the file can be edited locally after approval.
  • Provenance — the repository, marketplace or colleague it came from.

Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks.

Which mistakes do security teams make most often when vetting agent Skills?

The mistakes that security teams repeat when vetting Agent Skills cluster into a short, recognizable set. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — often just a markdown file on a laptop — and it executes a file read, an API call or a shell command with the employee's own permissions. Each row below pairs the corrective action with the tradeoff it carries.

Recurring mistake Risk consequence Do this instead — but watch out for
Treating the Skill description as the behavior contract The stated purpose and the executed commands diverge with nothing flagging it Review the instructions and scripts themselves; expect the file to change after review
Reviewing only the Skills an employee declares Undeclared Skills stay outside the inventory entirely, which is shadow AI by definition Discover continuously from the endpoint; a point-in-time scan goes stale quickly
Ignoring the identity the Skill runs under Actions inherit a developer's tokens, repository access and cloud credentials Map effective permissions, not nominal ones; personal accounts on corporate machines break the mapping
Assuming marketplace or vendor provenance equals safety A trusted publisher still ships an updatable file nobody re-reads Score components on observed capability; provenance stays useful signal, not a verdict
Reviewing only at first install Post-approval edits to the Skill, its hooks or its rules file land unreviewed Re-assess on change; noisy change events need policy tuning per team
Never recording what the Skill actually did No path from prompt to tool call to outcome when an investigation starts Retain tracing data; retention scope needs agreeing with privacy owners first

What if approval has already been granted?

Approved components still change after review. Backslash Security assesses the risk posture of endpoint agentic components on an ongoing basis and enforces allowlist and denylist policy as the agent runs, so an approval decision is re-applied rather than assumed.

Why does a one-time approval at install break down once a Skill starts acting?

A one-time approval at install captures what a Skill looked like on the day it was added, and little about how it behaves on its hundredth run. An Agent Skill — a packaged set of instructions and scripts that extends what an AI agent can do, frequently just a markdown file sitting on the machine — executes with the employee's own permissions, so an install-time review reads the text while the authority behind it goes unexamined.

Four things move after that review closes:

  • Remote context. The Skill pulls in files, repositories, issues or web pages the reviewer never saw, and that content becomes instructions the agent acts on.
  • Tool chaining. A Skill can reach an MCP server — a Model Context Protocol connection an agent acts through — inheriting reach the original review never scoped.
  • Upstream updates. The source changes; the local copy follows without a second look.
  • Standing credentials. Cloud keys, tokens and Git configuration already on the endpoint are available to whatever the agent does next.

This means approved-at-install says very little about approved-at-execution. Prompt injection — hostile instructions hidden in content the agent reads — shows how quickly capability expands: Backslash security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.

Do this But watch out for Mitigation in the same step
Approve Skills centrally The approved artifact changes upstream Re-assess the component continuously, not only at onboarding
Allow MCP connections per team Capability expands silently through chaining Scope which servers and tools each Skill may reach
Trust the author The content the Skill reads is untrusted Treat retrieved context as input, never as instruction
Log agent activity Logs arrive after the credential call Apply the control at the point of execution

How do review-time controls and execution-time controls compare for agent Skills?

Review-time controls and execution-time controls answer different questions about the same Agent Skill — a packaged set of instructions and scripts, often just a markdown file, that an agent can invoke with the user's own permissions. Review-time controls inspect that Skill before anyone is allowed to use it. Execution-time controls watch what the agent actually does when the Skill runs, and can stop the action before it completes. The two layers are complementary: approval decides what is permitted to exist on the endpoint, enforcement decides what is permitted to happen.

Before comparing them, it helps to fix the criteria that make the comparison decisive:

  • Visibility scope — what each layer can actually observe. Decisive when a Skill's behavior depends on files, repositories or MCP servers it reaches only at runtime.
  • Stopping power — whether the layer can prevent an action or only record a decision. Decisive in environments where an agent holds live credentials.
  • Failure mode — how the control degrades when it is wrong. Decisive for teams sizing residual risk.
  • Evidence produced — what an auditor or incident responder can reconstruct afterward.
  • Employee friction — the cost imposed on the people adopting AI, which governs whether the control survives contact with engineering.
Dimension Review-time control Execution-time control
What it sees Skill contents, manifest, declared permissions, source and author The live agent run: prompt, tool call, target, credential reach
What it can stop Introduction of a known-bad or unapproved component A specific action mid-run, before it executes
Failure mode Approved Skill later modified, or behaves differently against new context Policy too narrow blocks legitimate work; too broad permits drift
Evidence produced Approval record and version reviewed Trace of prompt, agent, tool call and outcome
Employee friction Queue and wait before use Minimal until a policy boundary is reached

Should Skills be vetted the same way as MCP servers, connectors, and rules files?

Skills deserve the same seriousness in review as MCP servers, connectors, and rules files, but they are not vetted by the same questions — each component sits at a different point in the agentic layer and goes wrong in a different way.

Start with the word itself, because "Skill" circulates with two meanings. In the catalog sense, an Agent Skill is a reusable capability an agent can invoke — an entry in a list of things the assistant can do. In the on-disk sense, it is packaged instructions and scripts, in practice often just a markdown file on the machine, that execute with the user's own permissions. This section uses the on-disk sense, because that is the artifact a reviewer can open, approve, or deny.

Component What it is What it carries What to ask before approval
Agent Skill Packaged instructions and scripts the agent invokes Instructions and intent, running with the user's permissions What does it instruct the agent to do, and what does it execute?
MCP server A connection to external tools and data through the Model Context Protocol Tools and reach into systems beyond the endpoint Which tools does it expose, and what can they touch?
Connector A link between the agent and an external destination Data egress Where does information go, and is that destination approved?
Rules file (AGENTS.md, CLAUDE.md) Standing instructions an agent reads on every run Silent, persistent shaping of behavior Who can edit it, and what does it say today?
Hook A trigger that runs something before or after an agent action Execution outside the prompt path What fires it, and what does it run?

Backslash Security scores MCP servers and Skills for risk and also judges their combinations: a Skill can be safe and an MCP server can be safe while the connection between the two is severely problematic.

What evidence should a Skill assessment actually collect before approval?

A defensible Skill assessment collects evidence about behavior and reach, not a source-code vulnerability review. The scope here is deliberately narrow: approving a single Agent Skill — a packaged set of instructions and scripts that extends what an AI agent can do, often nothing more than a markdown file on an employee's machine, executing with that employee's own permissions. The reviewer is judging an agentic component and the actions it enables, not auditing application code for flaws.

Evidence to collect Question the reviewer should be able to answer
Inventory across all employees and endpoints Where is this Skill already running, and on whose machines, before we formally approve it?
Inherited identity and permissions Which identity does it execute under, and is that a personal account or a managed one?
Reachable tools and external destinations Which MCP servers, APIs and outbound endpoints can it call once invoked?
Provenance and update history Who authored it, where did it come from, and can it change after approval without re-review?
Data classes present in its context What source trees, credentials, tickets or documents enter the context window during a run?
Reconstructable action record If this Skill misbehaves next quarter, can we replay prompt, agent, tool call and outcome?

A Skill's exposure follows the identity it borrows rather than the text it contains, which is why permission mapping belongs in the evidence pack alongside the file itself. Latio's 2026 AI Security Market Report found that "where Backslash stood out in our evaluation is in control depth," citing in-depth control over endpoint agent settings and over what is available to an agent in the first place. As of 2026, Backslash Security pairs that assessment data with allowlisting and denylisting, turning an approval decision into an enforced policy rather than a spreadsheet entry.

Frequently Asked Questions

What do security teams most often get wrong when vetting AI agent Skills?

The most common mistake in agent skills security is reviewing a Skill as if it were documentation. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — in practice often just a markdown file sitting on an employee's machine — and a Skill can tell an agent to read a file, call an API, or run a shell command, executing with that user's own permissions. Teams also tend to review at the moment of approval and never again, so the file that was read on day one is not the file running a month later. Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks, which gives a review process a concrete starting artifact.

Why is a point-in-time approval not enough for Skills, rules files, and hooks?

A Skill approved once keeps running alongside two components that reviews frequently skip: a rules file — a file in a repository or on a machine carrying standing instructions an agent reads on every run, such as AGENTS.md or CLAUDE.md — and a hook, a trigger that fires on an agent action and runs something before or after it. All three are editable text on the endpoint. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — an example of indirect prompt injection, where hidden instructions in content an agent reads are followed as though a user had typed them. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

How does vetting a Skill differ from assessing an MCP server?

A Skill packages instructions and scripts for the agent itself, while MCP — the Model Context Protocol — is how agents connect to external tools and data sources, with each MCP server acting as a connection an agent can act through. MCP server security review therefore asks what a connection reaches and what it is permitted to do, while Skill review asks what the agent has been told to do and with which permissions. Scale differs too: Backslash Security's MCP Server Security Hub, a public and continuously updated risk database, held 81,021 MCP servers when read on 22 September 2026, each scored for risk, according to Backslash Security. Both components sit on the same endpoint and interact, which is why assessing one in isolation leaves the other unexamined.

How can a team build an inventory without installing software on every endpoint first?

Start agentless — collecting information without leaving software permanently installed on the endpoint. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, distributed through existing device management tooling such as Intune or Jamf. That first pass is also where shadow AI surfaces: unapproved and invisible AI tools, agents, MCP servers and Skills running inside the organization, frequently installed by employees under personal accounts on corporate machines. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which converts an unknown population into something a governance program can report against.

Which AI agents should a Skills vetting program cover?

Coverage should follow what employees actually run, and as of 2026 that list is broad across both coding and general-purpose assistants. Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity. Latio's 2026 AI Security Market Report concludes that Backslash is "one of the most complete options available," with coverage spanning the full agentic surface from Claude Code and Cursor to the MCP servers and Skills running on top of them. A program scoped to a single approved assistant will miss the others that arrived without a procurement cycle.

What evidence should a Skills review leave behind for auditors and investigators?

A review should produce a record that survives the incident, not only the decision. That means a traceable path for each agent run and mapped audit output for the regimes the organization answers to. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and Backslash Security states that it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-08

Ready to make the switch?

See why teams choose Backslash Security.

Book a Demo