Blog

Vetting AI Agent Skills Before Approval: What Enterprise Security Teams Should Check

At a glance

  • Vet a Skill by reading what it instructs the agent to do and which files, credentials and connections it can reach.
  • An Agent Skill is often just a markdown file on an employee's machine, yet it executes with that employee's own access.
  • Vetting starts with discovery: unapproved Skills installed under personal accounts are shadow AI, surfaced by an endpoint inventory of agents, Skills and MCP servers.
  • Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader for runtime controls over MCPs and Skills on the endpoint.
  • Approval is continuous: allowlist approved components, denylist risky ones, block violating actions inline before execution, and retain audit records of agent runs.

Backslash Security

Published:

Security teams at AI-heavy enterprises — the organizations where a substantial share of the workforce actively builds or runs AI agents — should vet an AI agent Skill the way they would vet an unsigned script that carries full user privileges. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do; in practice it is often a single markdown file sitting on an employee's laptop, and it executes with that employee's own permissions, able to read files, call APIs or run shell commands. A workable approval sequence therefore runs: discover what is already installed across endpoints, read the Skill's actual instructions and the permissions it inherits, map the external connections it depends on, then record an allow or deny decision that is enforced continuously instead of once at intake.

Discovery comes first because few security organizations have any list of the Skills, agents, hooks and rules files running on company machines. Skills installed under personal accounts and never reviewed are shadow AI: AI tooling inside the company that security has not approved and cannot see. Review scope also extends past the Skill file itself, because Skills act through MCP servers — connections built on the Model Context Protocol that let an agent reach external tools and data — so MCP server security belongs inside the same approval question. Backslash Security's MCP Server Security Hub held 81,021 MCP servers, each scored for risk, when read on 22 September 2026, giving reviewers risk context for the connections a candidate Skill relies on.

What exactly is an AI agent Skill, and what does it let an agent do?

To say exactly what an AI agent Skill is: it is a packaged folder of instructions, scripts and supporting resources that an agent loads in order to extend its own behavior. In practice, the entry point is often just a markdown file sitting on an employee's machine, written in plain language, telling the agent how and when to perform a task. Because the Skill executes with the user's own permissions, it can tell an agent to read a file, call an API, or run a shell command without any separate authorization step.

The word carries two meanings that are easy to conflate. There is the assistant-marketplace sense — published extensions for consumer voice assistants, submitted to a vendor catalog, reviewed by that vendor and executed in the vendor's own cloud rather than on a corporate laptop. Then there is the agentic-endpoint sense, the canonical form in this context: Agent Skills loaded locally by coding and desktop agents such as Claude Code or Cursor, authored or downloaded by the employee, never reviewed by anyone. This article uses the agentic-endpoint sense throughout.

Skills are also routinely confused with their neighbors in the same layer:

Component What it is What it gives the agent
Agent Skill Packaged instructions plus scripts the agent loads A reusable capability it can invoke on demand
MCP server A connection built on the Model Context Protocol, the protocol agents use to reach external tools and data A live channel to act through
MCP tool An individual operation exposed by that server A specific callable action
Hook A trigger that fires on an agent action, running something before or after it Automatic execution around agent activity
Rules file (AGENTS.md, CLAUDE.md) A file carrying standing instructions read on every run Persistent behavioral defaults
Plugin / connector A packaged integration with another application Reach into that application

A Skill takes effect as soon as its file lands in the directory the agent reads: there is no installer, no package registry and no approval gate in between.

Why do Skills slip past the approval processes security teams already run?

When an employee adds an Agent Skill to their own laptop, it slips past the approval gates security teams already run, for a simple reason: it never reaches them. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — a Skill can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions. In practice, a Skill is often just a markdown file in a user directory. No installer, no admin rights, no purchase order, no vendor contract — so procurement review, SaaS security review, and the application inventory in Intune or Jamf have nothing to register.

Nothing here has to be malicious for the exposure to exist. EDR — endpoint detection and response, the incumbent layer built to catch malicious processes and known-bad files — observes a sanctioned binary such as Claude Code or Cursor behaving normally. What that binary was instructed to do, and which tool calls followed, sits above the process boundary it watches.

Which Skill attributes decide whether existing controls can see it?

Attribute Typical values Why it matters for approval
Artifact form Markdown instructions, optional helper scripts No executable to hash, sign, or allowlist
Install path User-writable directory, no elevation Bypasses software deployment and package inventory
Identity The employee's own session, tokens, and access Actions look like the user, not like an agent
Distribution Repository clone, copy-paste, marketplace, chat message No vendor to review, no contract to assess
Mutability Editable at any moment, no pinned version Approval of a snapshot does not bind later edits
Reach Local files, APIs, connected MCP servers Blast radius extends past the machine

Which risks should a Skill review actually be looking for?

This section narrows to a single artifact: one Agent Skill — packaged instructions and scripts that extend what an AI agent can do, often distributed as little more than a markdown file — and the risks a reviewer should enumerate before approving it. Each class below pairs a review action with the tradeoff that action leaves open, because a Skill executes with the permissions of the employee who installed it.

Do this in the review But watch out for this
Read every instruction line in the Skill, including sections the installer skims Standing instructions redirect agent behavior on every subsequent run; require a diff and re-approval on update rather than a one-time sign-off
Inspect bundled resources — templates, sample files, documentation shipped alongside the Skill — for prompt injection, meaning hidden instructions in content the agent reads and follows as if typed by the user Review only covers what exists at approval time; content the agent fetches later can carry the same hidden instructions, so pair the review with controls that act at execution
Enumerate the filesystem paths and network destinations the Skill touches Tight scoping breaks legitimate work; set reach per team and risk profile instead of one global rule
Check for reads of environment variables, token stores, cloud configuration and Git configuration Backslash security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — a path no manifest declares
Identify where the Skill sends data and under which identity Outbound movement is often indirect, routed through a tool the Skill calls rather than a direct request
Trace which MCP servers the Skill can invoke, MCP being the Model Context Protocol agents use to reach external tools and data Composition creates reach that neither the Skill nor the server shows alone; score the chain, not the parts
Verify provenance, pin the reviewed version, and record the hash Auto-update can silently replace approved content with unreviewed content under the same name

Each row describes a property a reviewer can establish by reading the artifact itself, and a residual exposure that only execution-time controls on the endpoint can close.

How should a security team structure a Skill vetting workflow step by step?

A security team can structure Skill vetting as a single intake-to-approval pipeline with named owners at each gate, rather than handling requests ad hoc. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — often little more than a markdown file on an employee's machine — and they execute with that employee's own permissions. If a Skill inherits the user's access, this means the review scope is everything that user can reach: repositories, cloud credentials, ticketing systems and connected MCP servers.

Teams at the point of deciding what to approve generally work through the following gates:

  1. Discover what is already installed. Inventory the Skills, agents, MCP servers, hooks and rules files present on endpoints today, so the review starts from a real population rather than from submissions people volunteer.
  2. Check provenance and authorship. Record where the Skill came from, who published it, whether the source is a maintained repository or a copied snippet, and who inside the organization sponsors it.
  3. Read the instruction layer. Open the markdown itself. Standing instructions, embedded prompts and references to external content are the part a reviewer can actually read before anything executes.
  4. Map declared capability against actual reach. Compare what the Skill says it does with the shell commands, file paths, APIs and tool calls it can invoke.
  5. Observe it in a constrained environment. Run the Skill against non-production data with narrowed credentials and watch which tools it calls and which destinations it contacts.
  6. Assign a risk tier. Tier by blast radius — read-only helpers, Skills that write to source control, Skills that touch secrets or production.
  7. Make the approval decision explicit. Allowlist approved Skills, denylist rejected ones, and attach the tier to a policy that can be enforced per team.
  8. Re-review on update. A Skill is a mutable file; treat any change to its instructions, scripts or declared tools as a new submission.

Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which gives reviewers at the observation and re-review gates a record of what an approved Skill did in practice.

What criteria belong on a Skill approval checklist, and how should each be scored?

The criteria that belong on a Skill approval checklist are the ones a reviewer can inspect statically, before the Skill ever executes. Agent Skills are packaged instructions and scripts that extend what an AI agent can do—often just a markdown file on an employee's machine—and they run with that employee's own permissions and access. That last property is why the rubric below weights reach and recoverability, not polish.

Define the scoring bands before you open the first file: approve (evidence present, no open questions), conditional (approve with a narrowed policy, a pinned version, or monitoring in place), reject (missing evidence or an irreversible capability nobody can justify). Each criterion is decisive in a different situation—provenance decides fast-moving consumer-grade Skills, permission scope and egress decide anything touching source repositories or production credentials.

Criterion What to inspect Evidence supporting approval Triggers rejection or conditional
Provenance Publisher identity, repository history, distribution channel Named, verifiable maintainer; Skill obtained from a controlled channel Anonymous or re-uploaded source; no history
Instruction transparency Full text of the markdown instructions and any embedded directives Instructions readable end to end and matching the stated purpose Obfuscated text, hidden directives, or instructions that read like a second prompt
Permission scope File paths, shell commands, APIs and tools the Skill invokes Scope narrow and explicitly bounded to the task Shell execution, credential paths, or unbounded filesystem reach
Data egress Outbound destinations, including any MCP server the Skill routes through Destinations enumerated and already approved Undeclared endpoints; personal-account connections
Dependency and update behavior Pinning, auto-update logic, transitive packages Pinned version; changes re-enter review Silent self-update—conditional at best, with change monitoring
Reversibility What the Skill writes, deletes, or pushes Actions observable and undoable Destructive or irrecoverable writes

A Skill's risk tracks the breadth of what it can reach rather than the sophistication of what it contains. Backslash Security carries the rubric's result forward into enforcement, allowlisting approved components and denylisting risky ones under custom policies per team. Record the approval against the pinned version rather than the Skill name: once it updates, the artifact a reviewer cleared is no longer the artifact running on the endpoint.

Frequently Asked Questions

What is an AI agent Skill, and why does it need vetting before approval?

An AI agent Skill is a packaged set of instructions and scripts that extends what an agent can do — read a file, call an API, run a shell command — and it executes with the user's own permissions. In practice a Skill is often just a markdown file sitting on an employee's machine, outside anything MDM or a software inventory tracks. Vetting matters because approval grants the Skill the identity and access of the person running it, not a sandboxed service account.

How can a security team assess a Skill without installing anything on the endpoint?

Assessing a Skill without permanent software on the machine is the agentless approach: collecting information without leaving an installed component behind. Backslash Security offers a free AI Endpoint Exposure Assessment — known in conversation as Scout — which is agentless, read-only, and retains no data, and Backslash Security also operates a free Skills Security Scanner that scans AI agent Skills for security risks. Agentless applies to discovery and assessment; enforcement on the endpoint is a separate control.

Which checks should a Skill review actually cover?

A practical Skill review covers the components the Skill can reach and the instructions it carries:

  • Declared capabilities — file reads, shell commands, outbound network calls, and credential access.
  • Connected MCP servers — MCP, the Model Context Protocol, is how agents connect to external tools, so each server is a path the Skill can act through. Backslash Security's MCP Server Security Hub, a public risk database, held 81,021 MCP servers, each scored for risk, when read on 22 September 2026.
  • Rules files — a rules file such as AGENTS.md or CLAUDE.md carries standing instructions an agent reads on every run, and Backslash security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.
  • Provenance and ownership — who published the Skill and whether it entered through a sanctioned channel.

How do teams catch Skills that were never submitted for approval?

Skills that were never submitted are shadow AI: agents, MCP servers and Skills running inside the organization that security has not approved and cannot see, frequently installed under personal accounts on corporate machines. Discovery has to run continuously rather than at intake, because the inventory changes whenever an employee installs a new extension. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, across tools such as Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity.

What controls apply after a Skill is approved?

Approval is the start of the control, not the end of it, because an approved Skill can still be manipulated into acting outside its task — a rogue agent. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. Latio's 2026 AI Security Market Report named Backslash an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills.

How do you evidence Skill approvals to auditors?

Evidencing Skill approvals requires a record of what was running, what was permitted, and what was blocked. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and according to Backslash Security it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. As Philip Walsh, Head of Security Engineering at Happy Returns, put it: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-08

Ready to get started?

See how Backslash Security can help.

Book a Demo