Blog

How to Scan an AI Skill for Risk Before Developers Install It: A Pre-Install Review Guide for Security Teams at AI-Heavy Engineering Organizations

At a glance

  • Treat an Agent Skill as executable code: review its instructions, scripts, commands, and reachable credentials before any developer installs it on an endpoint.
  • Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks.
  • Pre-install review covers only the moment of approval; Skills can be edited afterward, so continuous discovery and inline blocking still matter.
  • Per Backslash Security, intent-based behavior analysis detects and stops hundreds of potentially risky behaviors agents exhibit from poor goals or malicious intervention.

Backslash Security

Published:

To scan an AI Skill for risk before developers install it, treat the Skill as executable code rather than documentation: open the instruction file itself, enumerate every shell command, API call, and file path it invokes, establish which credentials, repositories, and directories it can reach under the developer's own permissions, and verify where it came from and how it updates. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — and in practice it is often just a markdown file sitting on a laptop, running with whatever access that employee already holds. For security teams at engineering-heavy organizations where a substantial part of the workforce builds or runs AI agents, that review has to happen at the point of approval and then keep running, because a Skill approved on Monday can be edited on Tuesday without any package manager, ticket, or device management console noticing.

Backslash Security, founded by Shahar Man (CEO and Co-founder) and Yossi Pik (CTO and Co-founder), operates a free Skills Security Scanner that scans AI agent Skills for security risks, and secures the broader set of components a Skill executes alongside: the agents themselves, MCP servers — connections an agent acts through under the Model Context Protocol — plus hooks, rules files, plugins, and connectors on the same machine. The guidance that follows reflects the Skill, MCP server, and rules-file formats developers are actually running on endpoints as of 2026, and it covers both halves of the problem: what to check before installation, and what has to stay in place once the Skill is live.

What is an AI Skill, and what does installing one actually add to a machine?

An AI Skill is a packaged unit of instructions and scripts that extends what an agent can do, and installing one on an endpoint is less dramatic than the word implies: in practice it is often just a markdown file, sometimes with a few helper scripts, dropped into a directory the agent reads. This section narrows to that single component — not the agent, the model, or the MCP server (Model Context Protocol is the protocol agents use to reach external tools and data), but the Skill file itself and what it puts on the machine.

Attribute What it contains Why it matters before approval
Artifact Markdown instructions, optional scripts or config Nothing compiles or registers with the operating system, so standard software inventory rarely records it
Execution context The user's own session and permissions The Skill acts with the employee's access to repositories, cloud profiles and internal systems
Declared actions Read a file, call an API, run a shell command Shell and network actions open the path to unauthorized execution and credential access
Dependencies MCP servers, connectors and plugins it calls through Each dependency is another route the agent can act through
Invocation Triggered when a task matches the Skill's description It can run without the user explicitly asking for it
Provenance Public marketplace, repository, teammate or vendor Unrecorded origin is how unsanctioned AI components spread between machines
Mutability A plain file on disk, editable after review What was approved and what actually runs can diverge

Those attributes define the scope of a pre-install review: open the file, read the standing instructions it gives the agent, list the commands and endpoints it declares, and follow every MCP server or connector it reaches through before the Skill lands in a developer's working directory.

Which risk signals should a reviewer look for inside a Skill before it is installed?

A reviewer reading an Agent Skill before installation is hunting a specific set of risk signals inside the package itself. This section narrows to that one task: the pre-install read of a single Skill. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — often little more than a markdown file plus a helper script — and it runs with the installing user's own permissions.

What to check in the Skill But watch out for — and how to handle it
Shell and process execution. Instructions telling the agent to run a command, invoke an interpreter, or spawn a child process. A Skill can describe the command in prose rather than code, so a text search misses it. Read the instruction body, not only the scripts.
Credential and config paths. References to cloud credential files, package registry tokens, SSH keys, or Git configuration. Legitimate Skills touch these too. Require the author to justify each path against the Skill's stated purpose.
Network destinations. Outbound URLs, webhooks, or runtime fetches of further instructions. If a Skill pulls its own content at run time, the version reviewed is not the version that executes. Pin it or reject it.
File scope. Directory globs reaching outside the project the agent is working in. Broad globs usually look like convenience. Narrow them before approval rather than denying the whole Skill.
Embedded directives to the model. Text telling the agent to suppress output, skip confirmation, or ignore prior rules. This is where indirect prompt injection lives: hidden instructions in content the agent reads, followed as if typed by the user. Treat any bypass of confirmation as disqualifying.
Declared MCP servers and connectors. Model Context Protocol servers are the connections an agent acts through, and a Skill may register one. The server, not the Skill, holds the real reach. Review the server separately before approving the Skill that calls it.

How do you verify where a Skill came from and who maintains it?

To verify where a Skill came from, split the question into three checks that the word "origin" usually blurs. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — often a single markdown file on a developer's machine — and it executes with that employee's own permissions. Origin can mean the distribution channel, the author, or whoever still maintains it, and each is confirmed differently.

Where did the file physically come from? Trace the install path: a marketplace listing, a public repository clone, a shared internal folder, or a paste from a chat thread. Pin an exact commit or tagged release rather than a branch, so the artifact reviewed today is the artifact that runs tomorrow.

Who published it? Prefer an organization-owned account over a personal one, check whether commits are signed, and confirm the publisher identity matches the name claimed in the documentation. A plausible README is not an identity.

Is anyone maintaining it? Look for recent commit activity, an issue tracker with answered issues, a changelog, and a stated security contact.

Trust signals worth recording before approval:

Signal What to look for Why it matters
Publisher identity Organization account, signed commits Ties the artifact to an accountable party
Version pinning Commit SHA or tagged release Prevents silent changes after review
Maintenance activity Recent commits, answered issues Indicates someone fixes defects
Independent assessment External risk assessment results Adds evidence beyond the publisher's own claims

Backslash Security assesses the risk posture of endpoint agentic AI components using an agentless approach — gathering that information without leaving software permanently installed on the machine.

What does a repeatable pre-install Skill review look like, step by step?

A repeatable pre-install review turns Skill approval from a judgment call into a checklist any reviewer can run the same way twice. Agent Skills are packaged instructions and scripts that extend what an AI agent can do — often little more than a markdown file plus a helper script — and they execute with the installing user's own permissions. This is decision-stage work: the reviewer already accepts that the risk exists and needs a defensible yes or no before the component lands on a workstation.

What are the steps in the review?

  1. Collect the full artifact. Pull the instruction file, any bundled scripts, and the manifest. A Skill is rarely a single file, so review everything that ships with it.
  2. Read the prose as code. Every instruction line is something the agent will act on. Flag directives to read credentials, write outside the project directory, or call an external endpoint.
  3. Enumerate the reach. List the shell commands, file paths, network destinations, and any MCP servers — Model Context Protocol connections an agent acts through — that the Skill can invoke.
  4. Check provenance and update behavior. Identify the publisher and source repository, then establish whether the Skill updates itself after approval. Self-updating content invalidates a one-time review.
  5. Run an automated pass. A consistent tooling baseline before manual judgment stops two reviewers from reaching different conclusions on the same artifact.
  6. Record the decision. Write the verdict into an allowlist or denylist entry carrying an owner, the approved version, and the date the record was made.
  7. Set a re-review trigger. Tie the approval to the reviewed version, so a new release, a changed rules file — a file carrying standing instructions the agent reads on every run — or an added hook sends the Skill back through the same sequence.

Which review layers catch which kinds of Skill risk?

Each review layer catches a different class of Skill risk, so the practical question is which layer catches what, and what slips past it to the one behind. An Agent Skill is a packaged set of instructions and scripts that extends what an AI agent can do — often just a markdown file on an employee's machine that tells the agent to read a file, call an API, or run a shell command with that user's own permissions.

Four criteria separate the layers, and each turns decisive in a different situation:

  • Timing — before install, at install, or at the moment an action executes. Decisive when a Skill is edited after approval.
  • Coverage — every Skill on every endpoint, or only the ones someone remembered to submit.
  • Behavioral visibility — what the Skill made the agent do, not only what its text declares.
  • Evidence — whether the layer leaves a record usable for incident reconstruction.
Layer Timing Coverage Behavioral visibility Evidence
Manual review Pre-install, one Skill at a time Only what is submitted Reader's judgment of static text A ticket, if one is filed
Automated assessment Pre-install and on discovery Every discovered Skill, MCP server and rules file Static and configuration risk signals Structured inventory and risk posture
Inline enforcement At execution Every agent action on the endpoint Direct — the action itself Trace of prompt, agent, tool call, outcome

A Skill's risk lives in the permissions it executes with rather than in the text an approver read last quarter. Manual review fits a short list of high-blast-radius Skills; automated assessment carries the inventory; inline enforcement handles whatever changes after approval. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

Frequently Asked Questions

What is an Agent Skill, and why does it need review before installation?

Agent Skills are packaged instructions and scripts that extend what an AI agent can do. A Skill can tell an agent to read a file, call an API, or run a shell command, and it executes with the user's own permissions — the same access the developer has to source repositories, cloud credentials and internal systems. In practice a Skill is often just a markdown file on the developer's machine, which means it arrives with none of the review gates that governed traditional software installs. Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks.

How do you check a Skill for risk before a developer installs it?

A workable pre-install review has four concrete parts:

  1. Read what the Skill actually instructs. Identify every file read, API call and shell command it can trigger, and the credentials those actions would touch.
  2. Scan it. Run the Skill through a dedicated Skills scanner, such as the free Skills Security Scanner Backslash Security operates, rather than relying on a visual read of the file.
  3. Assess what it connects to. A Skill frequently depends on MCP servers — Model Context Protocol endpoints that let an agent reach external tools and data. According to Backslash Security, its public MCP Server Security Hub held 81,021 MCP servers, each scored for risk, when read on 22 September 2026.
  4. Record a decision and enforce it. Allowlist approved components, denylist risky ones, and express the result as a policy that applies to the team in question instead of a one-time ticket.

Why doesn't our existing endpoint tooling flag a risky Skill?

Endpoint detection and response — the incumbent endpoint security layer — is built to catch malicious processes, files and known-bad behavior on a machine. A Skill is a set of standing instructions that a trusted agent binary reads and follows, so the activity it produces looks like ordinary, signed software doing ordinary work. Mobile device management governs the applications an employee installs; it has no view of the Skills, hooks, rules files and connectors an agent picks up afterward. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.

Can a pre-install review alone keep a Skill safe?

No — review establishes what a component is at a point in time, while agent behavior is shaped at run time by whatever content the agent reads. Prompt injection is the mechanism to understand here: hidden instructions placed in a file, repository, issue or web page that the agent follows as though the user had typed them. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

How do we find the Skills already running, and reconstruct what one did?

Start with the free AI Endpoint Exposure Assessment from Backslash Security — called Scout — which is agentless, read-only, retains no data, and is distributed as a short script through the device management system you already run. It surfaces shadow AI: the agents, MCP servers and Skills nobody approved, frequently installed under personal accounts. For the forensic half, Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. As Philip Walsh, Head of Security Engineering at Happy Returns, put it: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."

Who built Backslash Security, and has the approach been independently assessed?

Backslash Security was founded by Shahar Man, CEO and Co-founder, and Yossi Pik, CTO and Co-founder. On independent assessment, Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills. The same Latio report states that the company "was one of the first companies to focus on securing AI on the endpoint, building AI endpoint controls well before this year's rush of newcomers."


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to get started?

See how Backslash Security can help.

Book a Demo