At a glance
- Vet each MCP server against a fixed sequence: publisher, exposed tools, requested permissions and credentials, data handling, update path, isolated test, then allowlist entry.
- Discovery precedes vetting: employees install MCP servers locally under their own identities, so maintain a live inventory of what already runs on endpoints.
- Backslash Security's free AI Endpoint Exposure Assessment is agentless, read-only, and retains no data, making it a low-friction way to see what is installed.
- Re-check approved servers whenever versions, tool sets, or permissions change, and enforce allow and deny decisions on the endpoint itself.
Backslash Security
Published:
To vet an MCP server before an employee installs it, work through a fixed sequence: identify the publisher and source repository, enumerate every tool the server exposes and the permissions, credentials and network destinations it needs, review how it handles data and how it ships updates, run it in a contained environment against test data, and only then add it by name to an approved allowlist that is enforced on the machine. MCP, the Model Context Protocol, is the protocol AI agents use to connect to external tools and data sources, so each MCP server is a standing connection an agent can act through — typically installed by the employee, on their own workstation, under their own identity and access. Because that installation usually happens locally and without a ticket, a review process only holds if it is paired with continuous discovery of the agents, MCP servers, Skills, rules files and hooks already running on employee endpoints.
Scale shapes how that review has to be built. Backslash Security's MCP Server Security Hub had scanned and scored more than 80,000 publicly available MCP servers as of 22 September 2026, each one a candidate an employee could install. A workable program therefore combines a documented intake checklist for new requests, allowlist and denylist decisions enforced at the endpoint, and a durable record of what was approved and why — Backslash Security automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, which gives that record a form auditors and incident responders can both use. The steps that follow set out a repeatable MCP server security review that an IT, security or platform team can run without blocking the people who need these tools.
What exactly are you vetting when you vet an MCP server?
What exactly you are vetting depends on what you mean by "an MCP server," because the phrase covers two distinct objects and a reviewer has to vet both. MCP, the Model Context Protocol, is the protocol agents use to connect to external tools and data sources, and each server is a connection an agent can act through. The first meaning is the artifact — a package pulled from a registry or repository and run on the employee's machine. The second is the configured connection — the entry written into a client such as Claude Code, Cursor, Claude Desktop or Gemini CLI, which decides what that server is allowed to touch. A review that clears the package but never opens the configuration leaves the server's actual reach unexamined.
The discrete objects to put in front of a reviewer:
| Object | What it is and what varies | Why it matters to the review |
|---|---|---|
| Manifest / client config entry | The server definition in a client config file: command, arguments, environment variables, server name | Determines what actually launches on the endpoint; a familiar name can point at arbitrary code |
| Tool definitions | Each callable function the server exposes, with its description and input schema | Tool descriptions are read by the model as instructions, so the wording is part of the attack surface |
| Transport | Local stdio process, or a remote HTTP/SSE endpoint | Local transport inherits the user's own permissions; remote transport sends context off the machine |
| Credentials and scope | API keys, OAuth tokens, cloud profiles, repository access the server inherits | Scope is usually the employee's own identity, not a service account, so the blast radius is their full access |
| Publisher and provenance | Author, repository history, release signing, update channel | An unsigned or auto-updating package can change behavior after approval |
| Install surface | Where it lands, what it writes, what it can execute | Establishes whether the server can modify other components the agent reads, such as rules files or hooks |
Which risk signals should a pre-install review check first?
Scope note: this covers one candidate server under evaluation before it reaches an employee's machine — not the wider endpoint inventory, and not post-install monitoring. Under the Model Context Protocol, an MCP server exposes external tools and data to an AI agent, so every approval creates a path the agent can act through with the user's own permissions.
Six risk signals carry a pre-install review:
| Check | Do this | But watch out for |
|---|---|---|
| Publisher provenance | Verify the publishing organization, repository history, and release signing before approving. | A credible publisher name is easy to imitate on a package registry — match the repository to the vendor's documented source, not to search ranking. |
| Declared tool permissions | Enumerate every tool exposed and the file, shell, and API access each one implies. | Declared scope is self-reported; treat the manifest as a claim to test in a sandbox, not as a control. |
| Outbound destinations | Record which hosts the connector contacts and require an explicit allowlist. | Legitimate telemetry and model endpoints resemble exfiltration paths — classify them up front so reviewers aren't desensitized to alerts later. |
| Credential and token scope | Confirm which secrets are read, and issue a scoped, revocable token instead of reusing standing employee credentials. | Narrow scoping breaks some workflows; pair it with a documented request path so users escalate rather than pasting a broader key. |
| Install and update mechanism | Pin a version and review how updates arrive, including auto-update behavior. | Pinning freezes security fixes too — set a re-review trigger on each pinned version. |
| Prompt-injection exposure | Read tool descriptions and bundled instruction files as attack surface: hidden instructions placed in content an agent reads get followed as though the user typed them. | Descriptions change between releases, so a clean read at approval does not cover the next version. |
Credential scope and outbound destinations deserve attention early because they define the blast radius if any later check proves wrong: a connector with a narrowly scoped token and a short destination allowlist fails smaller than one holding the employee's standing access.
Record a written outcome per server — approved, approved with scope limits, or denied — together with the date and version examined, so the decision can be reconstructed during an investigation.
How do you run the vetting process step by step?
Run the vetting process as a fixed, repeatable sequence rather than an ad-hoc favor, so every MCP server — the connection an agent acts through to reach external tools and data — is requested, judged, and recorded against a named owner. This workflow belongs to the decision stage: a requester already wants a specific server, and security has to return a defensible answer fast enough to be used.
- Capture the request in your existing service desk. Require the server name, publisher and source repository, transport, the credentials or scopes it needs, which agent will load it (Claude Code, Cursor, GitHub Copilot), and the business purpose. Expected outcome: a ticket you can reopen on any future version change.
- Triage into a risk tier by blast radius. Separate read-only servers touching public data from those holding tokens to Git, cloud consoles, ticketing, or production. Expected outcome: each request routed to either a light review or a deep one, with the tier written on the ticket.
- Check publisher provenance and declared tools before installing anything. Read the exposed tool definitions, requested scopes, and update history, and weigh published assessments above the project's own description. Expected outcome: a documented risk rating and a list of the tools the server actually exposes.
- Trial the server under an isolated identity. Use a scoped, short-lived token and a non-production account on a contained workstation, then compare the tool calls it makes against the ones it advertises. Expected outcome: an observed behavior record, not a claimed capability list.
- Decide, then write the decision into enforcement. Approve to the allowlist, deny to the denylist, or grant a time-bounded exception with a named owner and an expiry date. Backslash Security carries that decision onto endpoints through custom policies tuned to different teams and risk profiles, enforced continuously. Expected outcome: the decision is live on managed machines rather than stored in a spreadsheet.
- Pin the version and set a re-review trigger. Treat a changed server build, a new tool definition, or an altered scope as a new request. Expected outcome: approvals expire on a defined trigger instead of aging silently.
Why does vetting at install time stop working once the server is running?
Vetting a Model Context Protocol server at install time produces a snapshot, and that snapshot stops describing reality the moment the server starts running. An MCP server — an endpoint an agent connects to and acts through — is not a static artifact like a signed binary. Its tool definitions are advertised dynamically when the agent connects, much of its logic executes remotely where a local reviewer has no visibility, and package updates can land without anyone reopening the approval ticket. This means an approval grants a capability that can change afterward, not a fixed piece of software.
The identity question compounds it. The agent connects using the employee's own tokens, keys, and sessions, so every call made through that connector is authorized as that person. Downstream systems in Okta, Entra ID, or a Git host record legitimate-looking activity, and the difference between sanctioned work and a manipulated agent lives in the intent behind the action.
| Do this | But watch out for | Handle it by |
|---|---|---|
| Re-enumerate tool definitions on every connection, not only at approval | Benign churn creates noise and reviewers start rubber-stamping | Alert only on changes adding permissions, shell commands, or new network destinations |
| Track the installed version of each connector | Remote behavior can change with no local version change | Watch what the service actually does, not what it reports itself to be |
| Let agents reuse existing corporate credentials | Misuse looks identical to normal work in downstream logs | Control at the action level — credential reads, privilege escalation, egress — not the identity level |
| Record agent activity for later review | Process-level telemetry shows that a binary ran, not what the agent was instructed to do | Capture the prompt-to-tool-call path so an investigation can reconstruct intent |
Closing this gap takes three capabilities running continuously rather than once: discovery that re-detects servers, Skills, and rules files as they appear; action control applied at execution; and a retrievable record afterward. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is the record an investigation needs when the approved connector behaved in a way nobody sanctioned.
Which criteria separate an approved MCP server from a blocked one?
The criteria that separate an approved MCP server from a blocked one are properties of the connection itself, not of the publisher's reputation. An MCP server — a connector built on the Model Context Protocol that lets an agent reach external tools and data — should be scored against named criteria agreed before any review starts, so reviewers debate exposure rather than brand familiarity.
Define the criteria first:
- Publisher trust tier — who maintains the server and how the build is distributed. Decisive when a server is community-published and installable straight from a package registry.
- Data sensitivity reached — what the connection can read once attached. Decisive when the server touches source control, ticketing, customer records, or secrets stores.
- Permission breadth — how many tools the server exposes and whether scopes are least-privilege. Decisive when a single install grants shell, filesystem, and network access together.
- Reversibility of actions — whether what the server does can be undone. Decisive for anything that writes, deletes, merges, deploys, or sends outbound.
- Update control — whether the version employees run is pinned and reviewed. Decisive when the server auto-updates from an upstream you do not control.
| Criterion | Approve | Approve with conditions | Deny |
|---|---|---|---|
| Publisher trust tier | Vendor-maintained, verified build | Known community maintainer, pinned release | Anonymous or unverifiable origin |
| Data sensitivity reached | Public or low-value data | Sensitive data, scoped to one system | Secrets, credentials, production stores |
| Permission breadth | Read-only, narrow tool set | Write scoped to named resources | Shell plus filesystem plus network |
| Reversibility | Read or reversible writes | Irreversible writes behind approval | Irreversible external sends |
| Update control | Pinned, change-reviewed | Pinned with re-review on bump | Silent auto-update |
Approvals of this kind collapse faster from unreviewed version bumps than from a misjudged first assessment, because the decision is a snapshot and the server keeps shipping. That is why the disposition has to become an enforced control: Backslash Security turns these outcomes into policy, allowlisting approved components and denylisting risky ones, with separate policies for teams carrying different risk profiles.
Frequently Asked Questions
What is an MCP server, and why does it need vetting before an employee installs it?
MCP (Model Context Protocol) is the protocol AI agents use to connect to external tools and data sources, so each MCP server an employee installs is a live connection an agent can act through — reading files, calling APIs, touching internal systems, all with that employee's own identity and permissions. Vetting matters because the install usually happens locally, in a client such as Claude Code, Cursor or Claude Desktop, with no ticket and no review. As of 22 September 2026, Backslash Security reports that its MCP Server Security Hub has scanned and scored more than 80,000 publicly available MCP servers, which gives an indication of how large the installable pool facing employees has become.
How can a security team assess MCP servers when they sit on an employee's own machine?
Start with discovery rather than policy, because an approval workflow for MCP servers is unenforceable without an inventory of what is already installed. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data — agentless meaning information is collected without leaving software permanently installed on the endpoint — which can be distributed through existing MDM tooling such as Intune or Jamf. The output is the baseline a vetting process is built on: which servers exist, on which machines, under which accounts.
Why doesn't the existing endpoint security stack flag a risky MCP server?
EDR — endpoint detection and response, the incumbent layer built to catch malicious processes, files and known-bad behavior — operates at the process boundary. A legitimate agent binary making a legitimate network call to an MCP server looks like normal activity, because nothing malicious has executed; what changed is the set of instructions and tools the agent now has. Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills.
What else should a vetting checklist cover besides the server itself?
An MCP server rarely arrives alone. Review the surrounding components an agent reads on every run:
- Agent Skills — packaged instructions and scripts that extend what an agent can do, often just a markdown file on the machine, executing with the user's own permissions.
- Rules files such as AGENTS.md or CLAUDE.md — files carrying standing instructions an agent reads automatically.
- Hooks, plugins and connectors — triggers and extensions that fire around agent actions.
Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — an example of prompt injection, where hidden instructions in content an agent reads are followed as though the user typed them.
Which records do auditors expect for approved and rejected MCP servers?
Auditors expect a defensible record of what was permitted, what was blocked, and what actually ran. Backslash Security automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, and traces the full path of an agent run from prompt to agent to tool call to outcome, which is the record an investigation needs after an incident. Approval decisions alone are not evidence; the execution trail is.
What happens when an approved MCP server starts behaving badly?
An approved component acting outside its task is a rogue agent — reaching for credentials, escalating privilege, or sending data to an unapproved destination — which is a different problem from Shadow AI, the unapproved tools, agents and Skills employees run that security cannot see. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. As Philip Walsh, Head of Security Engineering at Happy Returns, put it: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-08