At a glance
- A 30-day vetting plan runs in three stages: discover what employees already run, decide what is approved, then enforce those decisions on the endpoint.
- Plugin vetting now covers MCP servers, Agent Skills, hooks, rules files and connectors — not just browser extensions or IDE add-ons.
- Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, giving you a first inventory.
- Approval decisions only hold when they are enforced where the agent executes, on the host, rather than recorded in a spreadsheet.
- Treat the first cycle as a baseline, then repeat discovery continuously as employees install new components.
Backslash Security
Published:
You can stand up a working AI plugin vetting process in 30 days by splitting the month into three blocks: days 1 to 10 for discovery, days 11 to 20 for policy and an approval path, and days 21 to 30 for enforcement and audit trails. Start with discovery, because you cannot vet a catalog you have never seen — and in most organizations nobody has yet inventoried the Agent Skills, MCP servers, hooks, rules files, plugins and connectors sitting on employee laptops. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data, which makes it a practical way to produce that first inventory inside the opening block without waiting on a procurement cycle.
A short definition before the plan, because the word "plugin" has quietly expanded. In an agentic context it no longer means a browser extension or an IDE add-on. It means any component that extends what an AI agent can do: an MCP server, where the Model Context Protocol is the protocol agents use to connect to external tools and data sources and each server is a live connection an agent can act through; an Agent Skill, which is packaged instructions and scripts — often literally a markdown file on a developer's machine — that executes with the user's own permissions; a hook, a trigger that fires before or after an agent action; and a rules file such as AGENTS.md or CLAUDE.md, a file carrying standing instructions the agent reads on every run. Each of these is installed by an employee, under that employee's identity, with that employee's access. That is what makes vetting them an endpoint problem rather than a procurement one.
The 30-day shape below assumes you will not get clean results on the first pass. Treat month one as a baseline — a named inventory, a first approved list, a documented approval route, and a record of what agents did — and then repeat it, because the set of components on your endpoints changes every week. As of 2026, Backslash Security's research into malicious repository configurations and poisoned rules files is a reminder of why the standing instructions an agent reads deserve the same scrutiny as the tools it calls.
What exactly needs vetting before an AI plugin runs on an employee endpoint?
Knowing exactly what needs vetting starts with naming the artifacts, because an AI plugin is rarely one thing. This section narrows to a single case: one extension, Skill, or connector about to run on one employee's machine, under that employee's identity. Everything below is an attribute a vetting program records before approval, not after an incident.
| Artifact | What to record | Why it matters |
|---|---|---|
| Plugin or extension manifest | Publisher, version, install path, host agent (Claude Code, Cursor, GitHub Copilot, Codex) | Identifies what is actually present versus what was requested |
| Requested permissions | Filesystem scope, shell execution, network egress, credential store access | Permissions, not intent, set the ceiling on damage |
| MCP servers and tools | Server origin, transport, each exposed tool and its parameters | Model Context Protocol is how an agent reaches external systems; every server is a live path out |
| Agent Skills | Instructions and scripts the Skill contains, and the commands it can invoke | A Skill is often a markdown file that runs with the user's own permissions |
| Hooks | The agent action that fires the trigger and what runs before or after it | Execution can occur without a fresh user prompt |
| Rules files (AGENTS.md, CLAUDE.md) | Standing instructions the agent reads on every run, and who can edit them | Instruction content is loaded automatically at each invocation |
| Identity and data reach | Account type (corporate directory versus a personal login), repositories, tickets, and production systems in range | Determines blast radius and whether activity is attributable |
Two attributes deserve explicit values rather than a yes/no. Identity should be recorded as corporate — for example an account federated through Entra ID or Okta — or personal, since a private account on a managed laptop breaks attribution. Data reach should be recorded as the named systems in scope: Git remotes, cloud credentials, ticketing in Jira or ServiceNow, and any production endpoint.
Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which is what turns this list from a questionnaire into a populated record.
How do allowlisting, manual review, and inline action control compare as vetting approaches?
Allowlisting, manual review boards, policy-at-install gating, and inline pre-execution action control each vet AI plugins at a different layer, and they fail in different ways. Before comparing them, it helps to fix the criteria that decide which one fits a given risk.
- Coverage — what share of the agentic surface the control actually sees: agents, MCP servers (the protocol agents use to reach external tools and data), Agent Skills (packaged instructions and scripts that run with the user's own permissions), hooks, and rules files. Coverage is decisive when components arrive faster than anyone can catalog them.
- Latency — the delay between an employee wanting a component and being able to use it. This matters most where engineering velocity is a business constraint.
- Evidence — whether the approach leaves a reconstructable record of what was approved, what ran, and what it touched. Decisive for audit and incident investigation.
- Failure mode — what happens when the control is wrong or out of date.
| Approach | Coverage | Latency | Evidence produced | Failure mode |
|---|---|---|---|---|
| Static allowlist | Known components only; silent on anything new | Low once published, high for new requests | A list, not a runtime record | Goes stale; unlisted components run unobserved |
| Manual review board | Deep per item, narrow in volume | High — queue-bound | Rich written rationale per decision | Backlog pushes users around the process |
| Policy-at-install gating | Components present at install time | Low at install, none afterward | Install-time decision log | Blind to what an approved component does later |
| Inline pre-execution action control | The action itself, at the host agentic layer | Evaluated at the moment of execution | Trace of prompt, agent, tool call and outcome | Requires on-host enforcement to apply at all |
These layers stack rather than substitute. An allowlist settles what may be present; a review board settles the hard cases; install gating holds the boundary at onboarding; action control governs what an approved component does at the moment it acts. Backslash Security enforces at that agentic layer on the host, where the agent, Skill, or MCP server actually executes, and resolves which component caused a given action.
Days 1 to 10: how do you discover which AI plugins are already installed?
Days 1 to 10 exist to discover what AI is already running, not to judge it. The goal of this first phase is a complete, dated inventory of the agentic components installed on employee endpoints — and the identity each one executes under. Treat this as awareness-stage work: you are establishing facts, not issuing policy or revoking anything yet.
What are the concrete discovery steps?
- Define the endpoint scope. Pull the managed device list from Intune, Jamf or your existing MDM and agree which machines are in the first pass. All employees, all endpoints — developer workstations are the densest pocket, not the boundary.
- Enumerate agent clients. Look for Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, Antigravity, Devin and locally hosted models run through tools such as Ollama or LM Studio.
- Map the connection layer. Catalog every MCP server and MCP tool present. MCP, the Model Context Protocol, is how an agent reaches external tools and data, so each server is a live path an agent can act through.
- Collect the instruction artifacts. Record Agent Skills — packaged instructions and scripts that run with the user's own permissions — plus hooks, rules files such as AGENTS.md and CLAUDE.md that feed standing instructions into every run, and any plugins and connectors.
- Resolve identity. For each component, determine whether it authenticates through Entra ID, Okta or Active Directory, or through a personal account a developer signed in with. Personal logins on corporate machines are the signal that matters most to report upward.
- Record a baseline and stamp it. Write down the component counts alongside the date they were read, so the next pass measures change rather than restating a stale number.
Backslash Security's discovery covers each of these component types on the endpoint, including telling corporate accounts from personal ones, which is what the dated baseline in step 6 records.
Days 11 to 20: how do you turn that inventory into policy and an approval path?
Days 11 to 20 turn the raw inventory collected in the first stretch into written policy, named owners, and an approval path employees can actually follow. This is the decision stage of the program: the discovery question is answered, and the work now is classification and sign-off, not further exploration.
- Risk-tier what you found. Sort each discovered component — agents, models, MCP servers (Model Context Protocol connections an agent acts through), Agent Skills, hooks, plugins and rules files — into tiers by what it can reach. A Skill is often just a markdown file that tells an agent to read a file, call an API or run a shell command with the user's own permissions, so its tier follows its blast radius, not its size.
- Set the allowlist and denylist. An allowlist names approved components; a denylist names ones that may not run. Backslash Security supports both, plus custom policies per team and risk profile, so a research group and a finance team need not share one rule set.
- Assign an owner per tier. Give each tier a named accountable owner and a review cadence. Without one, exceptions accumulate in a ticket queue with no decision authority behind them.
- Write the exception path. Document who requests, who approves, what evidence is required, and how long an exception lasts. Route it through the tooling your teams already use for change approval.
- Decide what is blocked and what is only observed. Start by observing broadly and reserving denial for components with credential, repository or production reach, then tighten as the inventory stabilizes.
For the vetting evidence behind tiering decisions, Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks, which gives reviewers something concrete to reason about rather than a name and a version string.
Record every tiering decision with its rationale as you go. That record becomes the input for the enforcement and audit work in the final ten days.
Days 21 to 30: how do you enforce decisions and reconstruct what a plugin actually did?
In the final ten days you enforce the decisions reached during the approval stage and put in place the record that lets you reconstruct what a plugin, Skill or MCP server actually did. Enforcement is where a vetting program stops being a spreadsheet: the approved list becomes an allowlist, the rejected list becomes a denylist, and both are applied on the host where the agent executes rather than recorded after the fact. Backslash Security applies those guardrails at the agentic layer on the endpoint, which is also where the activity record is produced.
Day 21-24 — Convert the register into policy. Translate the approved and rejected components from the review stage into allowlist and denylist rules, with separate policy sets for teams whose risk profiles genuinely differ.
Day 25-27 — Pilot enforcement with one team. Run the policy against a single engineering group first, in a mode that surfaces what would have been stopped, before widening it.
Day 28-29 — Plan for approved components that act beyond their task. An approved agent can still take an action nobody asked for, whether after a prompt injection or simply while trying to complete its task. Inventory controls will not catch that, so decide which actions — credential access, privilege escalation, data sent to unapproved destinations — your policy blocks inline regardless of which approved component requests them.
Day 30 — Prove you can reconstruct a run. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is the record an incident reviewer needs when a component behaves unexpectedly.
| Do this | Watch out for — and how to handle it |
|---|---|
| Enforce denylists on the host | Blunt blocks interrupt legitimate work; pilot in observation mode first and tune per team |
| Allowlist approved components | New Skills and MCP servers appear between reviews; keep discovery running continuously rather than quarterly |
| Add action-level policy | Approved components can still take actions nobody asked for; pair the allowlist with rules about what an agent may do, not only what it is |
| Retain run-level traces | Traces are only useful if someone reads them; name an owner for post-incident review before day 30 |
Why do existing endpoint and governance layers leave AI plugin vetting uncovered?
When AI agents land on employee machines, the existing endpoint and governance layers an organization already owns keep doing exactly what they were built to do — and none of them was designed to vet what an agent loads.
Start with the word itself, because two different things get called a plugin. One is the marketplace extension inside a sanctioned SaaS application, governed through that application's admin console and the identity provider in front of it. The other is a component of the agentic AI fabric on an employee's own machine: an MCP server (a connection an agent acts through under the Model Context Protocol), an Agent Skill (packaged instructions and scripts that run with the user's own permissions, often just a markdown file), a hook, a connector, or a rules file. This article means the second.
What each layer was not built to see:
- EDR — endpoint detection and response watches processes, files and known-bad behavior. The coding agent is signed and approved, and its process tree looks identical whether the user typed the instruction or an agent read it somewhere.
- MDM — device management governs operating system configuration and installed applications. A Skill or rules file is a text file in a user directory; a connector is a configuration entry. Neither registers as an installation.
- SaaS and identity governance — covers sanctioned applications and corporate accounts, not an agent acting under an employee's personal login on a corporate machine.
- Network-layer inspection — resolves destinations and traffic, not which agent, Skill or MCP server initiated the call.
The gap here is one of causality rather than coverage: every layer records what executed, while the instruction that caused it sits in content the agent read on its own.
Backslash Security's own research illustrates the consequence — malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. A rules file carries standing instructions an agent reads on every run, and editing its contents changes nothing in the device-management record.
Frequently Asked Questions
What counts as an "AI plugin" when you scope a 30-day vetting program?
Treat "plugin" as shorthand for the whole extension layer an agent can load on an endpoint: MCP servers (Model Context Protocol connections an agent acts through), Agent Skills (packaged instructions and scripts that run with the user's own permissions — often just a markdown file), hooks (triggers that fire before or after an agent action), rules files such as AGENTS.md or CLAUDE.md that carry standing instructions on every run, plus connectors and models. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which is what makes a single scoped inventory possible.
How do you build the day-one inventory without deploying software first?
Backslash Security offers a free AI Endpoint Exposure Assessment — agentless, read-only, and retaining no data — distributed through your existing MDM such as Intune or Jamf. Backslash Security calls this assessment Scout. It gives you a list to vet against before any policy decision is made.
Why review rules files and hooks rather than only the plugin package?
Because the instruction surface is part of the attack surface. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — an indirect prompt injection, where hidden instructions in content the agent reads are followed as if a user typed them.
Which free resources can a team use while the program is still being built?
Backslash Security operates a free Skills Security Scanner that checks AI agent Skills for security risks, and publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks. According to Backslash Security, its MCP Server Security Hub, a public and continuously updated risk database, held 81,021 MCP servers when read on 22 September 2026, each scored for risk, giving reviewers reference scoring for servers already in use.
Can a 30-day effort produce evidence an auditor will accept?
Partly, and the mechanism matters more than the calendar. Per Backslash Security, the platform automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, and traces the full path of an agent run from prompt to agent to tool call to outcome. As of 2026, that trace is what lets a reviewer reconstruct which approved component took a given action.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07