At a glance
- A Skill review covers its instruction text, file and shell reach, the identity it runs under, its provenance, and whether its activity is recorded.
- Agent Skills execute with the employee's own permissions, so the review scope is what the Skill can reach, not only whether it looks malicious.
- Hidden instructions inside content an agent reads can redirect a trusted Skill, so provenance and update path matter beyond first install.
- Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks.
- Process-level endpoint telemetry records that a command ran, but not which Skill, agent or MCP server caused it.
Backslash Security
Published:
A security review of an AI agent Skill should cover what the Skill instructs the agent to do, what it can reach, and whose permissions it runs with. An Agent Skill is a package of instructions and scripts that extends what an AI agent can do — in practice, often just a markdown file sitting on an employee's machine — and it can tell an agent to read a file, call an API, or run a shell command, executing with that user's own account and access. The working scope of a review therefore includes the instruction text itself, the tools and MCP servers the Skill chains into (MCP, the Model Context Protocol, is how agents connect to external tools and data sources, and each MCP server is a connection the agent can act through), the credentials and directories reachable by the account running it, the file's origin and update path, and whether anything records what the Skill actually did afterward.
Two conditions make that review harder than a conventional software check. First, a Skill nobody approved is shadow AI: AI tooling running inside the organization that security has not sanctioned and cannot see, usually installed by an employee on their own endpoint, sometimes under a personal account. Second, endpoint detection and response tooling — the incumbent layer built to catch malicious processes and known-bad behavior on a machine — records the process that executed, not which prompt, agent, Skill or MCP server caused it to execute. As of 2026, Skills sit on employee endpoints alongside hooks, plugins and rules files such as AGENTS.md or CLAUDE.md, which carry standing instructions the agent reads on every run.
This is the layer Backslash Security was built for. It discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, assesses the risk posture of each, and blocks risky agent actions inline before execution — including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. Backslash Security is backed by Stage One, First Rays Venture Partners, Artofin and D E Shaw & Co, alongside a group of individual investors.
What does an AI agent Skill actually expose that a security review has to cover?
An AI agent Skill exposes considerably more than the capability its name advertises, and a security review has to treat every layer inside it as separately reachable at runtime. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — in practice often just a markdown file sitting on an employee's machine — and they execute with that employee's own permissions. This section narrows to that one component of the endpoint agentic layer; MCP servers, hooks and rules files carry their own review questions.
The reviewable attributes of a Skill:
- Instruction text. Free-form natural language the agent follows on invocation, unconstrained in content and rarely read end to end. This is where prompt injection lives: hidden instructions placed in content an agent reads, which it acts on as though the user had typed them.
- Bundled scripts and commands. Shell, Python or inline commands the Skill can trigger. Values range from read-only helpers to arbitrary execution, and they inherit the user's shell, not a sandbox.
- Declared tool and MCP calls. Each Model Context Protocol server — the protocol agents use to reach external tools and data sources — is a live connection the Skill can act through, so agent skills security and MCP server security are the same review in practice.
- File and path scope. Which directories the Skill reads or writes. Credential stores, Git configuration and environment files are the paths that matter most.
- Network destinations. Any address the Skill can send output to, approved or otherwise.
- Provenance and mutability. Where the Skill came from, who may edit it locally, and whether a change is visible to anyone. A markdown file can be rewritten after approval.
- Executing identity. The account, tokens and session the Skill runs under, which is usually the employee's own.
Backslash Security assesses the risk posture of endpoint agentic AI components using an agentless approach, so a reviewer works from what is installed on real machines.
Which checks belong on a Skill security review checklist?
The checks that belong on a Skill security review are concrete and enumerable, because an Agent Skill — packaged instructions and scripts that extend what an AI agent can do — is often just a markdown file on a machine, plus whatever that file is permitted to invoke. A Skill executes with the user's own permissions, so the review has to establish reach: what it can read, what it can run, and what it can connect to.
| Do this check | Risk it retires | But watch out for |
|---|---|---|
| Read every file the Skill ships, including nested instruction files | Hidden instructions the agent follows as if the user typed them | A one-time read goes stale the moment the Skill updates |
| Compare the declared purpose against actual shell commands and API calls | Capability creep beyond what the Skill advertises | Dynamically assembled commands a static read misses |
| Enumerate credential, token, and Git configuration paths it touches | Silent secrets access under the employee's identity | Legitimate Skills need some credentials; scope them instead of banning them |
| List network destinations and MCP servers it connects through | Data leaving for unapproved endpoints; weak MCP server security | An allowlist drifts as the agent's tool surface expands |
| Identify the hooks and rules files around it — files carrying standing instructions read on every run | Controls applied outside the Skill itself | Reviewers inspect the Skill and skip the rules file governing it |
| Verify publisher, source, and update channel | Untrusted components installed outside any approval path | Personal-account installs that never reach the managed inventory |
Does the review end at approval? No. The same file can be edited after sign-off, so agent skills security depends on continuous enforcement as much as on inspection. Backslash Security establishes guardrails for endpoint agentic AI — allowlisting approved components, denylisting risky ones, and enforcing custom policies per team continuously.
Who is accountable for a Skill nobody reviewed? Accountability sits with whoever owns AI risk, which in most organizations means the security team, even when the Skill was installed by an engineer under a personal account on their own workstation.
How do pre-install review, runtime observation, and inline action control compare as review layers?
Pre-install review, runtime observation, and inline action control each answer a different question about the same Agent Skill — a packaged set of instructions and scripts that extends what an AI agent can do, and that executes with the user's own permissions. Fix the evaluation criteria before comparing the layers:
- Coverage — which components the layer sees: the Skill file, the MCP server it calls through, the rules file shaping the run.
- Timing — whether assessment happens before installation, during execution, or afterward. Decisive when the action is irreversible, such as credential access.
- Evidence quality — what reconstructable record remains, which matters for audit obligations and investigation.
- Blind spots — what the layer structurally cannot observe, regardless of tuning.
| Layer | Coverage | Timing | Evidence produced | Structural blind spot |
|---|---|---|---|---|
| One-time pre-install review | The Skill's contents as submitted, plus declared dependencies | Before first use | An approval record fixed at a point in time | Later edits to the file, and behavior once live tool access is attached |
| After-the-fact runtime observation | Observed agent and tool activity across the session | During and after execution | Trace material for forensics | No action is prevented; the record is written after the outcome |
| Inline pre-execution action control | The action an agent is about to take, resolved to the Skill, agent or MCP server behind it | At execution, before completion | A decision log tied to the action and its cause | Intent that never reaches an enforceable action boundary |
Programs usually start at the first layer because it is cheapest to stand up, and a markdown file edited after approval is the common reason that record goes stale. Observation after the fact produces the trace, but the activity it describes has already run to completion. Backslash Security enforces on the host at the agentic layer, resolving which agent, Skill or MCP server is behind an action while the action is still pending.
Where each layer fits:
- Low agent volume and a slow rate of change: pre-install review carries most of the weight.
- Audit and forensics obligations: runtime observation supplies the reconstructable record.
- Irreversible actions on sensitive machines, including developer workstations: inline control is the layer that applies at execution time.
Why does Skill provenance and update behavior matter after approval?
A Skill's provenance — who authored it, which repository or marketplace it came from, and how it reaches the machine — and its update behavior both keep moving after approval is signed off. Agent Skills are packaged instructions and scripts that extend what an AI agent can do, often arriving as little more than a markdown file plus a few helper scripts, and they execute with the user's own permissions. This means approval attaches to a specific version of those files, not to the Skill's name: if the source refreshes itself, the artifact that runs tomorrow may share nothing with the one that was examined except its label.
So provenance has to be recorded as data, and update behavior has to be governed as a control surface in its own right.
| Do this | But watch out for — and how to contain it |
|---|---|
| Record author, source repository and distribution path at approval, and store them alongside the approved version | Maintainership can change hands quietly; bind approval to a verified publisher identity and re-assess when that identity changes |
| Pin the approved version and fingerprint the Skill's files on disk | Skills that pull fresh content from a remote source at run time sidestep the pin; monitor on-disk files for drift and re-review on change |
| Route installation through a reviewed internal channel | Employees can install straight from public sources under personal accounts; continuous discovery, rather than a one-time onboarding scan, is what surfaces that |
| Review the scripts and tool calls a Skill can invoke alongside its written instructions | New capability usually arrives in an update, not in the reviewed version; enforce at execution so an altered Skill still cannot exceed policy |
Backslash Security establishes guardrails for endpoint agentic AI through allowlisting of approved components and denylisting of risky ones, with custom policies for different teams and risk profiles enforced continuously rather than only at install time.
How do employee endpoints change what a Skill review can assume?
When an employee runs an Agent Skill on a machine they control, the endpoints themselves change what a review can safely assume. The term "Skill review" carries two distinct meanings, and they are not interchangeable.
The catalog review. This is an inspection of the published artifact: an Agent Skill — packaged instructions and scripts that extend what an AI agent can do — is read, approved, and added to a list. A Skill in practice is often just a markdown file, so the review means someone reads the file, checks what it tells the agent to do, and signs off. Example: an architect reviews a Skill in a shared repository before it reaches the approved set.
The installed-instance review. This asks what is actually present on a given machine right now, which agent loads it, and whose identity and permissions it executes with. A Skill can instruct an agent to read a file, call an API, or run a shell command, and it does so with the user's own access. Example: an employee copies a Skill from a public source into a local directory, or edits an approved one in place, and the approval path never sees either event.
This section uses the installed-instance meaning, because that is where a central process loses its grip. Mobile device management governs devices and applications, so a file an agent reads at runtime falls outside it. Endpoint detection and response resolves processes and binaries without resolving which prompt, agent, or Skill caused the action. A personal account signed in on a corporate machine compounds it.
Scope here spans all employees and all endpoints, with developer workstations one prominent use case inside it. Backslash Security enforces on the host at the agentic layer, where the agent actually executes.
What happens after the review, and who governs Skills in production?
What happens after a Skill security review is a cycle, not a filing: the verdict becomes an allowlist or denylist entry, that entry becomes enforceable policy, and the policy is applied on the endpoint every time the Skill runs. A Skill — packaged instructions and scripts that extend what an AI agent can do, executing with the user's own permissions — is on most machines little more than a markdown file in a folder, which gives a one-time verdict a short shelf life.
What should a team do first?
- Discover before you govern. Build the inventory of agents, MCP servers, Skills, hooks, rules files and plugins actually present on endpoints; Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, which keeps the first pass low-friction for IT.
- Assign a verdict per component, not per vendor — approved, restricted, or denied, with the reason recorded.
- Translate verdicts into policy, differentiated by team and risk profile, so a research group and a finance group need not share one rule set.
- Enforce where the agent executes. Backslash Security enforces on the host at the agentic layer, rather than relying on after-the-fact alerting.
- Keep the trace. Retain records linking the prompt to the tool call so an investigation can reconstruct what ran and under whose identity.
- Re-assess on change, treating any edit to an approved Skill or rules file as a new review trigger.
Review programs erode not at assessment but at re-approval, since an approved file edited the next morning still carries yesterday's verdict.
Ownership, as of 2026, commonly splits three ways: security sets policy, IT and MDM administrators maintain coverage across the fleet, and the AI governance lead reports against the inventory.
Frequently Asked Questions
What should a security review of an AI agent Skill cover at minimum?
A security review of an AI agent Skill should cover every capability the Skill grants and every resource it can reach. An Agent Skill is a package of instructions and scripts that extends what an AI agent can do — in practice often just a markdown file on a developer's machine — and it runs with that user's own permissions. Work through at least these items:
- Files and directories the Skill reads from or writes to
- Shell commands, scripts, or binaries it invokes
- Credentials, API tokens, and environment variables within its reach
- Outbound destinations it contacts, and whether they are approved
- MCP servers it calls, where MCP (Model Context Protocol) is the protocol agents use to connect to external tools and data sources
- Who authored it, where it came from, and how updates arrive
Why do rules files and hooks belong in the same review?
A Skill rarely acts alone on an endpoint. A rules file — AGENTS.md or CLAUDE.md, for example — is a file in a repository or on a machine carrying standing instructions the agent reads on every run, and a hook is a trigger that fires on an agent action, running something before or after it. Both can change what a reviewed Skill ends up doing. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens, and Git configuration — a Skill review that stops at the Skill file misses that path.
How can a team review Skills that nobody formally approved?
Start with discovery, because unreviewed components are the common case. Shadow AI means AI tools, agents, MCP servers, and Skills running inside the organization that security has not approved and cannot see — often installed by employees on their own endpoints under personal accounts. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin, and connector running on an endpoint, which converts an unknown population into a reviewable list. Backslash Security also offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, meaning nothing is left permanently installed on the machine.
Which free tools can check a Skill before it is approved?
Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks, which gives a reviewer a repeatable first pass instead of a manual read of each markdown file. For the related question of agent-level exposure, Backslash Security publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks. Reviewers assessing the connection side of the problem — MCP server security — should treat each MCP server as a live channel the agent can act through, not as a static dependency.
What record should a Skill review leave for auditors and incident response?
The review should produce a durable trace, because approval at a point in time does not explain what happened afterward. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is the record an investigator needs to reconstruct a Skill's behavior after the fact. According to Backslash Security, the platform automatically generates audit evidence for EU AI Act, NIS2, DORA, and SOC. Pair that trace with enforcement: as of 2026, Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
Who builds and backs Backslash Security?
Backslash Security was founded by Shahar Man, CEO and Co-founder, and Yossi Pik, CTO and Co-founder, to secure the agentic AI fabric — the mesh of agents, MCP servers, Skills, plugins, connectors, and hooks running on enterprise endpoints. Backslash Security is backed by Stage One, First Rays Venture Partners, Artofin, and D E Shaw & Co, alongside a group of individual investors. Its coverage includes AI coding agents such as Claude Code, Cursor, Codex, Devin, GitHub Copilot, and Antigravity.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07