At a glance
- Approve Agent Skills through continuous discovery and allowlisting on the endpoint, instead of routing every request through a ticket queue.
- A Skill is often just a markdown file, executing with the developer's own permissions on their own machine.
- Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data.
- Per-team policies let security approve broadly, deny narrowly, and keep engineering velocity intact while governing agent activity.
Backslash Security
Published:
Approving AI Skills at speed means replacing case-by-case review with a standing allowlist: discover every Skill already running on employee endpoints, assess its risk posture, approve the safe ones, deny the risky ones, and enforce that decision at the moment an agent acts. An Agent Skill is a package of instructions and scripts that extends what an AI agent can do — in practice often just a markdown file on a developer's machine — and it can tell the agent to read a file, call an API, or run a shell command with that developer's own permissions. Because the Skill executes on the endpoint under the user's identity, the approval decision has to be enforced on the endpoint too.
The incumbent control at that location is EDR, endpoint detection and response, bought to catch malicious processes, files, and known-bad behavior. Established platforms such as CrowdStrike and SentinelOne are purchased for exactly that job, with broad general-purpose coverage and an agent already installed. A Skill invocation, an MCP server call, or an instruction in a rules file is not a malicious process; it is an approved agent doing what something told it to do, which puts agent Skills security at a layer below the network and above the operating system. Backslash Security blocks risky agent actions inline before execution — unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations — and, as of 2026, covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity.
What exactly is an AI Skill, and why does it need an approval decision at all?
What exactly counts as an AI Skill depends on which layer of the agent stack you mean, because the term is used for at least three different things. Some teams mean the vendor feature by name — Agent Skills in products such as Claude Code or Claude Desktop. Others use it loosely for any packaged capability an agent can call, which sweeps in MCP tools, plugins and connectors. A third group uses it internally for a saved prompt template. This section uses the first meaning: Agent Skills are packaged instructions and scripts that extend what an AI agent can do. A Skill can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions.
Which attributes decide whether a Skill is approvable?
| Attribute | Range of values | Why it matters to the decision |
|---|---|---|
| Form on disk | Often a single markdown file; may bundle scripts | No installer and no package manager record, so mobile device management inventories do not list it |
| Execution identity | The employee's own account, tokens and session credentials | Reach equals the user's reach, including Git and cloud sessions |
| Invocation | Agent-selected at runtime, or user-invoked | A Skill can fire without a deliberate human action |
| Provenance | Internally authored, copied from a public repository, or shipped with the agent | Determines whether anyone has read what it instructs |
| Adjacent artifacts | Rules files such as AGENTS.md or CLAUDE.md, hooks, MCP servers, models, context | Components interact, so a Skill is not assessable alone |
A rules file here means a file on the machine or in a repository carrying standing instructions an agent reads on every run. Those standing instructions are part of the approval question: Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.
Why do Skill approvals slow developers down in the first place?
When Skill approvals run through a manual ticket queue, they are slow for a structural reason: the reviewer is working without the one thing that would make the review quick — a record of what is already running. An Agent Skill is a reusable capability an agent can invoke, and in practice it is often just a markdown file on a developer's machine that tells the agent to read a path, call an API, or run a shell command with that user's own permissions. Nothing in a standard MDM estate, such as Intune or Jamf, enumerates those files, so the ticket becomes the only control point.
Four mechanics produce most of the drag:
- No endpoint inventory. Reviewers cannot confirm whether a Skill is new, already in use elsewhere, or a renamed copy, so every request starts from zero.
- Opaque contents. A ticket shows a name and a repository link, not the file system paths, shell commands and MCP servers — the Model Context Protocol connections an agent acts through — that the Skill will actually reach.
- Duplicate review. The same popular Skill is re-reviewed by platform, data and application teams independently, filling the queue with repeat work.
- Mismatched clocks. Installing a Skill takes seconds; approval takes days, and developers close that gap by installing under their own identities.
That last mechanic produces shadow AI — AI tools, agents, MCP servers and Skills running inside the organization that security has not approved and cannot see, often under a personal login on a corporate machine.
| Do this | But watch out for | Handle it by |
|---|---|---|
| Stand up a formal Skill review board | Review latency pushing installs underground | Seeding the board with a live inventory; Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint |
| Block unapproved Skills outright | Developers switching to an agent you govern even less | Allowlisting approved components per team rather than banning the category |
| Delegate approval to each team | Inconsistent decisions on the same Skill | Custom policies per risk profile, enforced continuously against one shared record |
Which approval models can an enterprise actually choose between?
Enterprises choosing an approval model for Agent Skills are choosing between four architectures, and the differences surface in daily operations long before they appear in policy documents. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — often just a markdown file on a developer's machine — and they execute with that user's own permissions. Set the evaluation criteria before weighing any option:
- Time to approval — how long a developer waits between requesting a Skill and using it. Decisive where agent tooling changes week to week.
- Coverage of unapproved installs — whether the model sees a Skill that was never submitted for review. This is where shadow AI, meaning agent components security never approved and cannot see, either surfaces or stays hidden.
- Handling of Skill updates — whether approval survives the file changing after review, since a markdown file can be edited silently.
- Developer friction — how much of the cost lands on engineers rather than on the security function.
- Forensic reconstruction — whether you can later show what a Skill did, which prompt triggered it, and which tool call followed.
| Approval model | Time to approval | Unapproved installs | Skill updates | Developer friction | Forensic reconstruction |
|---|---|---|---|---|---|
| Manual ticket review | Days; queue-bound | Invisible — only submissions are seen | Re-review needed, rarely triggered | High | Ticket history only |
| Static allowlist/denylist of sources | Fast once a source is listed | Partial — covers listed sources | Source stays approved as content changes | Low after setup | Policy records only |
| Periodic endpoint audit | Not an approval path | Found at the next audit cycle | Caught only between cycles | Low | Point-in-time snapshots |
| Continuous discovery with policy enforcement at execution | Immediate for approved components | Discovered as they appear | Assessed on discovery | Low; decisions applied at execution time | Run-level traces |
Manual review suits a small, slow-moving set of high-blast-radius Skills. Source allowlisting fits organizations standardizing on a short list of internal repositories. Periodic audits serve compliance reporting where no approval decision is required. Continuous discovery with enforcement fits environments where engineers add components faster than any queue can process them — the layer Backslash Security operates at, and where, by Backslash Security's own account, audit evidence for EU AI Act, NIS2, DORA and SOC is generated automatically.
What should a Skill risk assessment actually look at before approval?
A Skill risk assessment narrows the question to a single component: the Agent Skill itself — packaged instructions and scripts that extend what an AI agent can do, often arriving as little more than a markdown file on a developer's machine and executing with that user's own permissions. Every attribute below is machine-readable, so approval can run as an automated check against policy instead of a queue waiting on a human reviewer.
| Attribute | What to capture | Why it decides approval |
|---|---|---|
| Origin and publisher | Source repository or registry, publisher identity, whether it was authored in-house or pulled from a public listing | An unattributed publisher leaves no one accountable and no signal to re-check when the source changes |
| Invocable tools and MCP servers | Which MCP servers — the Model Context Protocol connections an agent acts through — and which tool calls the Skill can reach | Defines the real blast radius; a modest Skill wired to a broad connector inherits that connector's reach |
| Filesystem and network reach | Paths it reads or writes, shell commands it may run, outbound destinations | Determines whether the Skill can move repository contents or local files off the machine |
| Credential and identity usage | Environment variables, token stores, cloud profiles, Git configuration it touches | Tokens reached under a personal identity extend access to systems the Skill was never scoped for |
| Instruction and prompt content | The literal standing text the agent will follow, including anything pulled in at runtime | Hidden instructions placed in content an agent reads are followed as though the user typed them |
| Update channel | How new versions arrive, and whether updates are pinned, signed, or silent | An approved version only holds if the next one cannot replace it unreviewed |
| Chaining | Hooks, subagents, rules files or other Skills it invokes | Components interact, so a Skill approved alone can compose into behavior nobody assessed |
Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which supplies the inventory these attributes attach to. Re-evaluate each attribute whenever a Skill's source or version changes, so the record stays current between formal reviews.
How can security teams approve Skills in minutes instead of weeks?
Security teams can approve Agent Skills in minutes rather than weeks by moving the control point away from the review queue and into continuous policy enforcement. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — often little more than a markdown file on a developer's machine — and it executes with that developer's own permissions.
The operating model has four moving parts, and they run in sequence:
- Discover first, review second. Maintain a continuous inventory of the Skills, MCP servers, hooks and rules files already present on endpoints. Organizations routinely find components in use that were never submitted for review, which is where the backlog originates.
- Write policy once, per risk profile. Define allowlists of approved components and denylists of risky ones, then vary the thresholds by team — platform engineering and a finance analytics group need different limits, not different processes.
- Assess automatically on arrival. Each newly observed Skill is evaluated against that policy as it appears, rather than when someone remembers to file a ticket. Backslash Security assesses the risk posture of endpoint agentic components using an agentless approach, so assessment keeps pace with what employees install.
- Tier the decision. Components inside policy clear without human involvement. Only exceptions — an unknown publisher, a shell command, a credential path — route to a reviewer, arriving with the assessment evidence already attached.
Approval then stops being the single moment at which risk is controlled. Backslash Security keeps policy enforced on the host while agentic activity is running, so a component acting outside its approved remit is handled at that moment rather than relied upon to have been judged correctly weeks earlier. A reviewer who knows this is in place can afford a faster default.
A practical starting point for teams evaluating this model is an inventory of the Skills already live on a sample of engineering endpoints, taken before any policy rule is written.
How do you roll out a fast Skill approval program without a big-bang change?
Roll the program out in stages and it lands faster than a big-bang cutover, because each stage produces the evidence the next one depends on. Agent Skills — packaged instructions and scripts that extend what an agent can do, often just a markdown file sitting on a developer's machine — appear and disappear faster than any monthly review board meets, so sequencing is what keeps approvals current.
Start in observe-only mode. Backslash Security's discovery maps the agents, MCP servers, Skills, hooks, rules files and connectors present on each endpoint, which gives you a real inventory to write policy against instead of a survey of what people say they use.
| Stage | What to measure | Common failure mode |
|---|---|---|
| Observe-only discovery | Endpoint coverage; distinct Skills, MCP servers and rules files seen; share tied to personal rather than corporate accounts | Treating first-week volume as a crisis and freezing adoption before tiers exist |
| Publish policy and risk tiers | Share of discovered components that map cleanly to a tier | Tiers written around tools rather than the actions a Skill can take |
| Auto-approve the long tail | Share approved without human review; time from first sighting to decision | Routing every low-risk Skill through the same queue as a credential-touching one |
| Inline enforcement on high-risk action classes | Blocked actions per week; appeal volume; developer-reported friction | Enabling enforcement broadly on day one instead of on a narrow action set |
| Extend to all employees and endpoints | Non-engineering endpoint coverage; unapproved tooling reaching sanctioned systems | Leaving the program scoped to engineering, where it was easiest to pilot |
Where these programs stall is almost always at the moment the inventory grows past the review capacity that produced it, which is the argument for tiering before enforcement.
Developer workstations are the natural pilot because agentic activity is densest there, but the same Skills, connectors and personal-account logins surface across sales, finance and support endpoints once discovery runs company-wide.
Frequently Asked Questions
What is an Agent Skill, and why does it need an approval step at all?
An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — it can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions. In practice a Skill is often just a markdown file sitting on a developer's machine. Approval matters because the Skill inherits the engineer's access to source control, cloud credentials, and production tooling. Backslash Security also operates a free Skills Security Scanner that scans AI agent Skills for security risks, which gives reviewers a starting point before a Skill reaches a workstation.
How can a security team approve Skills quickly instead of becoming a queue?
The workable pattern is to replace ticket-by-ticket review with standing policy: allowlist the components that have already been assessed, denylist the ones that are known to be risky, and let everything in between route to a short review. Backslash Security supports exactly that model — allowlisting of approved components, denylisting of risky ones, and custom policies tuned to different teams and risk profiles, enforced continuously rather than at a single gate. Developers feel a decision, not a wait. As Philip Walsh, Head of Security Engineering at Happy Returns, put it: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."
What happens when someone installs a Skill or MCP server nobody approved?
That is shadow AI: agents, models, MCP servers and Skills running inside the organization that security has not approved and cannot see, frequently installed under a personal account on a corporate machine. MCP, the Model Context Protocol, is how agents connect to external tools and data, so each MCP server is a live path an agent can act through. Detecting unapproved components is a discovery problem first. Per Backslash Security, its public MCP Server Security Hub had scanned and scored more than 80,000 publicly available MCP servers when read on 22 September 2026, so a reviewer looking at an unfamiliar public server may find a published risk score to start from.
Why doesn't the endpoint security stack we already own cover approved Skills?
Endpoint detection and response (EDR) was built to catch malicious processes, files and known-bad behavior on a machine. The agentic AI fabric — the interconnected layer of agents, models, MCP servers, Skills, hooks, rules files, plugins and connectors — operates above that boundary, and an approved agent invoking a Skill produces activity that looks ordinary at the process level while the record carries no trace of which prompt, agent or Skill caused it. Backslash Security works at that agentic layer: it blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. Latio's 2026 AI Security Market Report named Backslash an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills.
How do we prove to an auditor that our Skill approvals actually held?
Audit evidence has to show both the policy and the behavior under it. Backslash Security collects and presents tracing data of agentic activity, tracing the full path of an agent run from prompt to agent to tool call to outcome, and according to Backslash Security it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. That record is also what an incident investigation runs on: when a Skill or rules file — a file carrying standing instructions an agent reads on every run — produces an unexpected action, the trace shows which component issued it. Backslash security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration, which is the class of event this tracing is meant to reconstruct.
Where should a team start if it has no Skill inventory today?
Start with discovery before policy, because an allowlist written without an inventory just formalizes guesses. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless — meaning information is collected without leaving software permanently installed on the endpoint — read-only, and retains no data; it is distributed through existing MDM tooling such as Intune or Jamf. The output tells you which agents, MCP servers, Skills, hooks and rules files are already running, which is the input a first allowlist needs. From there, teams typically approve the components already in heavy use, denylist the ones that scored badly, and move the remainder into review.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07