Blog

How to Vet AI Skills Without Slowing Developer Velocity

At a glance

  • Vet Agent Skills where they run: continuous discovery, automated risk assessment, allowlisting of approved Skills, and inline blocking before a risky action executes.
  • A Skill is often just a markdown file on a laptop, executing with the developer's own permissions and never passing a review gate.
  • Backslash Security publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks.
  • A free AI Endpoint Exposure Assessment from Backslash Security is agentless, read-only, and retains no data.
  • Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader for in-depth endpoint AI controls.

Backslash Security

Published:

You can vet AI Agent Skills without adding a review queue by moving the check to where the Skill actually runs: discover every Skill installed on the endpoint, assess its risk automatically, allowlist the ones that clear, and block the risky action inline before it executes. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — a Skill can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions. In practice a Skill is frequently nothing more than a markdown file sitting on a developer's machine, which is why a manual approval step struggles here: engineers add and edit Skills faster than a ticket queue can clear them, and the file as written at approval time is not the same artifact as the actions the agent takes when it later runs.

Agent Skills security is one layer of agentic AI endpoint security — securing the agents, models, MCP servers, Skills, rules files and hooks that run on an employee's endpoint with that employee's identity and access. As of 2026 that layer spans mainstream developer tooling: Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity, together with the MCP servers — Model Context Protocol connections an agent acts through to reach external tools and data — installed alongside them. The numbered steps that follow set out a workflow a security architect, endpoint administrator or AI governance lead can run end to end: build the inventory, define what approved means for each team, enforce it at execution time, and retain the tracing data needed to reconstruct an agent run afterward.

How can engineering teams vet AI Skills without adding a review queue?

Engineering teams can vet Agent Skills at adoption speed by making the check automatic at installation, rather than standing up a human review queue. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do—often just a markdown file on the machine—and executes with the user's own permissions, which is why a file-level check matters. The workflow below keeps approval asynchronous.

Step 1 — Inventory the Skills already present on endpoints. Enumerate every Skill, along with rules files and hooks beside it, before writing any policy. Expected outcome: a named list of Skills per machine, with the commands and paths each one touches. Watch out: a point-in-time export ages quickly; bind discovery to a continuous mechanism so the list stays current.

Step 2 — Assess each Skill the moment it appears. Trigger assessment on first sight instead of on request, so the instructions, scripts and permissions a Skill carries are read automatically. Backslash Security performs this assessment of endpoint agentic components without leaving software permanently installed to collect it. Expected outcome: every newly installed Skill carries a risk rating without a developer filing a ticket. Watch out: a static assessment describes the file as written; behavior during an actual run needs its own control, so a clean result is not a permanent pass.

Step 3 — Allowlist cleared Skills and denylist the rest by policy. Express the decision as policy that varies by team and risk profile, so a platform group and a support group do not inherit the same permissions. Expected outcome: developers install from a pre-cleared set without waiting on security. Watch out: an over-broad allowlist silently re-creates the blind spot; scope entries to specific Skills, not whole categories.

Step 4 — Enforce at the action boundary. Apply guardrails that act on what the Skill attempts—credential reads, shell execution, outbound destinations—before the action completes. Expected outcome: risky actions are stopped while approved work proceeds untouched. Watch out: enforcement that only alerts produces a record of damage already done; configure prevention, not notification alone.

What is an AI Skill, and what makes one risky?

An AI Skill is a packaged set of instructions and scripts that extends agent capabilities, running with the employee's permissions on their machine. This section focuses on Skills on laptops or workstations invoked by agents like Claude Code, Cursor, or GitHub Copilot—not Skills on managed servers. In practice, a Skill is often just a markdown file in a folder, and security teams typically cannot see it at all.

A Skill can instruct an agent to read files, call APIs, or run shell commands. Risk comes from concrete attributes that can be inspected before approval:

Attribute What it can be Why it matters on an endpoint
Permission scope Inherited from the signed-in user; no separate identity The Skill can reach anything the employee can reach, including corporate SaaS sessions and local credentials
Invoked tools Shell commands, HTTP calls, MCP servers and MCP tools Each invoked tool is an outbound path the agent can act through, well beyond the text the Skill appears to contain
Context access Repository files, environment variables, configuration files, rules files such as AGENTS.md or CLAUDE.md Instructions hidden in content the agent reads can be followed as if the user typed them
Update behavior Edited locally, pulled from a public source, or overwritten silently between runs An approved version is not permanent; the file reviewed on Monday may not be the file that executes later
Provenance and trigger Author unknown or unverified; fires automatically on an agent action via a hook Nothing prompts a human to re-read the file at the moment it runs

Where do unvetted Skills enter employee and developer machines?

Unvetted Skills enter employee and developer machines through low-friction channels: copied install commands, cloned repositories, and plain text files users create without administrative rights. Agent Skills are packaged instructions extending AI agent capabilities—often just a markdown file in a folder the agent reads. Nothing crosses procurement, change tickets, or endpoint control.

Common arrival routes:

  • Registry and marketplace installs — MCP servers (Model Context Protocol for external tools and data) added directly to local agent configuration.
  • Cloned repositories — rules files like AGENTS.md or CLAUDE.md carrying standing instructions agents read on every run, plus hooks firing before or after agent actions.
  • Copy-paste from community content — Skills or plugins dropped into the agent's working directory.
  • Personal logins — employees connecting corporate machines to agents under private accounts, the most concrete shadow AI form.
  • Agent-initiated additions — agents pulling in connectors mid-run to finish tasks.
Do this But watch out for Mitigation in the same move
Publish an approved list of agents and MCP servers Users still add components locally, outside the catalog Pair the list with continuous discovery so additions surface as they appear
Require review for new Skills Review stalls work and pushes installs underground Keep review asynchronous and scoped to the component, not the task
Block personal account sign-ins Blunt blocks break legitimate workflows Detect the private-account connection specifically, rather than the tool

Does MDM already cover this?

Device management enrolls machines and manages applications and profiles. It does not read agent configuration directories. Backslash Security discovers those components on the endpoint itself—agents, MCP servers, Skills, hooks, rules files, plugins and connectors.

Which vetting approaches keep developer velocity intact?

Vetting approaches for Agent Skills — packaged instructions and scripts that extend AI agent capabilities, often just markdown files — differ sharply in preserving engineering throughput. Key criteria:

  • Speed to decision — time between wanting a Skill and running it. Decisive when agents are used daily and queues become workarounds.
  • Coverage — how much of the endpoint's agentic layer the method sees: Skills, MCP servers, hooks, rules files (AGENTS.md, CLAUDE.md), plugins, connectors. Decisive when components install under personal identities rather than MDM.
  • Operational cost — recurring human effort per decision. Decisive for small security teams supporting many agent builders.
Approach Speed to decision Coverage Operational cost
Manual security review Slow; gated on reviewer availability Deep on submissions, blind to unreported items High and linear with volume
Static allowlist Fast for listed items, blocking for new Catalog only; silent on unlisted/modified components Moderate, concentrated in upkeep
Periodic inventory audit No gate, no delay Broad but point-in-time; drift invisible between cycles Moderate, spiky per cycle
Inline pre-execution policy enforcement Immediate; decision at action Continuous across running components and behavior Low per decision once policy defined

Manual review suits few Skills with wide blast radius. Allowlists suit stable toolchains with rarely changing approved sets. Periodic audits suit reporting obligations requiring dated snapshots. Inline enforcement suits environments where Skills and MCP servers change between audits — the model Backslash Security applies: policy evaluates when agents attempt actions, so no review queue sits between a developer and an approved Skill.

What does end-to-end Skill governance look like across every endpoint?

End-to-end Skill governance is one stage in a continuous lifecycle across every endpoint, not a standalone approval gate. The parent discipline is agentic AI endpoint security: governing the agentic AI fabric—agents, models, MCP servers, MCP tools, Skills, hooks, rules files, plugins and connectors—that employees run on their machines under their own identities. An Agent Skill is packaged instructions and scripts extending agent capabilities; often just a markdown file executing with the user's permissions.

Stage What it answers Where Skill vetting fits
Discover Which agents, MCP servers, Skills, hooks and rules files exist on which machines? Produces the inventory a review queue draws from
Assess What risk posture does each component carry? Scores the declared actions, file reads and shell calls
Govern Which components are allowed, for which teams? Turns a passed review into an allowlist entry, a failed one into a denylist entry
Block inline What should be stopped before it executes? Enforces the approved boundary at run time
Reconstruct What actually happened, and in what order? Supplies forensic evidence when behavior surprises you

Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, making the reconstruct stage usable during investigations.

Skill review is a point-in-time artifact check, while the persisting risk is the permission the Skill borrows at execution, not the text approved weeks earlier.

Teams evaluating control sets in 2026 can check one thing: whether discovery, policy and enforcement draw on a single endpoint record, or hand off between separate inventory, policy engine and log pipeline.

Frequently Asked Questions

What counts as an AI Skill, and why does it need vetting?

An AI Skill is a packaged set of instructions and scripts that extends what an AI agent can do — it can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions. In practice a Skill is often just a markdown file sitting in a folder on a developer's machine, added in seconds and invisible to the security team. That combination is why agent skills security belongs in a vetting workflow: the artifact looks like documentation, but it behaves like executable capability inheriting a trusted identity.

How do you vet Skills without creating a review queue that stalls developers?

You vet Skills at the policy layer rather than ticket by ticket. Define an allowlist of approved components and a denylist of risky ones, write separate policies for teams with different risk profiles, and let anything inside policy run without human review. Backslash Security enforces those policies continuously and blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. Teams that want a starting point can run Skills through the free Skills Security Scanner that Backslash Security operates for assessing Skill security risks.

Which AI coding agents and components should be in scope?

Scope the review around the agents your engineers actually run plus everything attached to them. Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity, and discovery extends to every model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector on the endpoint. MCP, the Model Context Protocol, is how agents connect to external tools and data, so each MCP server is another path a Skill can act through. A rules file such as AGENTS.md or CLAUDE.md carries standing instructions the agent reads on every run, which makes it part of the same review set.

How do you find Skills nobody approved?

Start with discovery, because unapproved components are the defining case of shadow AI — AI tools, agents, MCP servers and Skills running inside the organization that security has not approved and cannot see, including components connected through a personal account on a corporate machine. Device management enrolls endpoints and process-level endpoint detection watches executables; neither enumerates what an agent can invoke. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, meaning it collects the inventory without leaving software permanently installed. It distributes through tooling such as Intune or Jamf, reports, and exits.

How do you assess risk in open-source agent projects like OpenClaw?

Treat an open-source agent project the way you treat any dependency: enumerate what it installs, what credentials it can reach, and what it is permitted to execute. For OpenClaw specifically, Backslash Security publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks, which gives teams a concrete starting inventory instead of a questionnaire. The same discipline applies to any agent runtime an engineer pulls down independently, since the risk lives in the permissions and connections it acquires on the machine, not in the project's popularity.

What audit evidence should you keep after a Skill is approved?

Keep a record that reconstructs behavior, not just approvals. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is what an investigator needs when an approved component starts acting outside its task — a rogue agent, as distinct from an unapproved one. Backslash Security also states that it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, which reduces the manual collection burden on security and compliance owners.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-08

Ready to get started?

See how Backslash Security can help.

Book a Demo