Comparison

How to Vet an MCP Server Before You Approve It Internally

At a glance

  • Vetting an MCP server means checking publisher provenance, the tools it exposes, credential scope, data destinations, and behavior once an agent connects.
  • Approval should end in an allowlist enforced continuously on the endpoint where the agent actually executes, under the employee's own identity.
  • Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks.
  • Backslash Security publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks.

Backslash Security

Published:

Vetting an MCP server before internal approval comes down to answering a defined set of questions about one piece of software: who publishes it, what tools it exposes to an agent, which credentials and scopes it asks for, where it sends data, and what it is allowed to do once it is connected. MCP, the Model Context Protocol, is the protocol agents use to reach external tools and data sources, which makes each MCP server a live connection an agent can act through — usually on an employee's own machine, under that employee's identity and access. That is what MCP server security means at the point of approval: a provenance and publisher check, a review of the exposed tool surface, an authentication and secrets review, a look at network destinations, and a judgment about runtime behavior. As of 22 September 2026, Backslash Security reports that its public MCP Server Security Hub held 81,021 MCP servers, each scored for risk, giving reviewers a reference point to check a candidate against. Once a server is approved, the decision still has to hold on the machine: Backslash Security discovers the agents, MCP servers and Skills running on employee endpoints, assesses which ones are safe to use, and blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

What exactly are you vetting when you review an MCP server?

Vetting an MCP server depends on what exactly you mean by the term, so a review has to separate two objects before it starts. MCP (Model Context Protocol) is the protocol agents use to connect to external tools and data sources, and each MCP server is a connection an agent can act through. In one sense, you are reviewing an artifact: a package, binary or configuration entry a user adds to Claude Code, Cursor or a desktop client. In the other, you are reviewing a running capability: a live service that exposes callable functions to an agent, usually under the employee's own identity and existing access. Approving one is closer to granting an agent standing authority to act on someone's behalf inside whatever that server reaches than it is to installing an application.

Which components make up the review surface?

Component What it covers Why it matters to the approval decision
Tools The callable functions the server exposes — read, write, execute, network calls Each tool is an action the agent can take; write and shell-capable tools carry the broadest blast radius
Prompts Server-supplied instruction templates the agent may load Instructions arriving from outside the user are a path for hidden directives in content the agent reads
Resources Files, repositories, databases or APIs the server can surface as context Defines what data can leave the machine and reach the model
Transport Local process communication versus a remote network endpoint Determines whether traffic leaves the endpoint and who operates the far side
Identity Tokens, OAuth grants or keys the server holds, and whose account they belong to Personal credentials on a corporate machine turn an approved connection into Shadow AI — unsanctioned AI use security cannot see
Update channel How new versions, tools or prompt definitions arrive after approval A server can gain capabilities after review without any further user action

Each row is a separate question for the reviewer, and a server can pass on five and still fail on the sixth.

Which review criteria separate an approvable MCP server from a risky one?

The review criteria that separate an approvable MCP server from a risky one come down to evidence: each criterion asks what the publisher, the package, or the running process can actually show you before an agent is allowed to act through it. MCP — the Model Context Protocol — is how agents connect to external tools and data sources, so every server is a live path an agent can take under the employee's own identity and access. Agree on the criteria, and on what counts as proof for each one, before any specific server is on the table.

Provenance and maintenance signals establish whether anyone is accountable for the code, and they are decisive when a server is community-published rather than vendor-released. Tool manifest and permission scope govern blast radius, and matter most when the server reaches a Git system, a cloud account, or a production environment. Authentication, token handling, and egress destinations determine where credentials and data can end up, which is decisive wherever the server brokers access to enterprise systems. Transport model and versioning behavior determine whether what you approved is what keeps running — the key question for remote servers that can change without a reinstall.

Criterion Evidence that supports approval Signal that warrants a hold
Publisher provenance Named publisher, signed release, source repository matching the published package Anonymous author, mirrored package, no link between repo and artifact
Tool manifest and permission scope Enumerated tools, each with a stated purpose and least-privilege scope Broad shell or filesystem tools, undeclared capabilities added at runtime
Authentication and token handling Short-lived tokens, scoped credentials, no secrets written to local config Long-lived keys in plaintext config, shared service accounts
Data egress destinations Documented endpoints, limited to the vendor's own domains Undisclosed third-party telemetry, dynamic destinations
Update and versioning behavior Pinned versions, published changelog, explicit upgrade step Silent auto-update, floating "latest" tag
Transport model Local transport for host-only tools; authenticated, encrypted remote transport Unauthenticated remote endpoints, mixed local and remote behavior
Maintenance signals Active issue handling, disclosed security contact Abandoned project, unresolved security reports

Backslash Security turns a completed scorecard into an enforced control, allowlisting the components a review has cleared and denylisting the ones it has not, with custom policies for different teams and risk profiles.

How do local, remote, and self-hosted MCP server deployments change the vetting questions?

Local, remote, and self-hosted MCP servers all raise the same review question — what can this connection reach, and under whose authority — but each architecture moves the answer somewhere different. An MCP server is a connector an AI agent acts through under the Model Context Protocol, exposing tools, files, or APIs to the agent. A local server runs as a process on the employee's own machine, usually over stdio (the standard input/output channel); a remote server is operated by a provider and reached over the network; a self-hosted server is the same code running on infrastructure your organization controls.

Define the criteria before comparing the options:

  • Identity — which account the server's actions resolve to. Decisive wherever an agent can act on production systems, because a personal token makes attribution impossible.
  • Data residency — where prompt content, file contents, and tool results come to rest. Decisive under contractual or regional data obligations.
  • Credential exposure — what secrets the server can read as a side effect of running. Decisive on developer workstations, where cloud keys, registry tokens, and Git configuration sit in predictable paths.
  • Observability — what record exists afterward. Decisive for incident reconstruction and for audit regimes such as the EU AI Act, NIS2, DORA, and SOC 2.
Deployment Identity Data residency Credential exposure Observability
Local (stdio) The employee's own OS user and permissions Stays on the endpoint unless a tool call sends it out Widest practical reach: local files, environment variables, keychains No network record; visible only on the host
Remote (provider-hosted) A token or OAuth grant, often issued to a personal account Prompt and tool data leaves your boundary Limited to granted scopes, which need their own review Provider-side logs you do not own
Self-hosted Your directory, if wired to Entra ID, Okta, or Active Directory Within your own infrastructure Depends on how the service account was scoped Yours to instrument, and yours to build

A local connector leaves no network trace, so approval depends on evidence collected from the host itself; Backslash Security enforces at the agentic layer on the endpoint, where a stdio server actually executes. For remote servers, read the requested scopes and confirm the grant is tied to a corporate identity in Entra ID or Okta. For self-hosted deployments, treat the service account scope and the log pipeline as part of the approval record.

What does a step-by-step MCP server approval workflow look like?

This workflow narrows the broad governance question down to a single object: one MCP server, one request, one decision. MCP — the Model Context Protocol — is how agents connect to external tools and data sources, so each server is a live connection an agent can act through, usually with the requesting employee's own identity and permissions. The step-by-step sequence below is written for the consideration-and-decision stage, where a security architect or endpoint administrator already accepts that approval is needed and wants the stages, owners, and exit criteria.

  1. Intake request. Capture the requester, the business task, the server's origin and publisher, the tools it exposes, and the credentials or scopes it will need. Reject nothing yet; the goal is a reviewable record.
  2. Discovery of what is already installed. Check the fleet before judging the request, because the same server may already be running somewhere unapproved. Endpoint-side discovery turns intake into a comparison against reality rather than a guess.
  3. Static manifest review. Read the configuration and tool definitions: declared tool descriptions, network destinations, filesystem paths, shell invocation, and any standing instructions the server injects into an agent's context.
  4. Sandboxed behavioral trial. Run it on an isolated host with test credentials and observe actual tool calls — what it reads, where it sends data, and whether behavior matches description. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is what makes a trial reviewable afterward.
  5. Risk tiering. Sort the result by blast radius: read-only and low-sensitivity, write access to internal systems, or access to source control, secrets, and production.
  6. Conditional approval with inline action controls. Approve the server with constraints attached to its tier, so that high-consequence actions — credential reads, privilege changes, outbound transfers to destinations outside the approved set — are stopped at the point of execution rather than reported later.
  7. Periodic re-review. Re-run discovery and re-check the manifest on a defined cadence and after any version change, since an updated server can add tools nobody reviewed.

Why do approved MCP servers still create risk after the review is signed off?

An approved MCP server can drift away from what was actually reviewed, because MCP servers — the connections an agent acts through under the Model Context Protocol — keep changing after sign-off. Approval attaches to a configuration at a single moment, and that configuration does not hold still. This means the control expires quietly: the allowlist entry remains, while the trust decision behind it no longer describes what the agent is talking to.

Four mechanisms account for most of that decay:

  • Silent tool definition changes. A server can revise a tool's description, parameters or behavior with no version signal the reviewer ever sees.
  • Prompt injection through returned context. Hidden instructions sitting in content the agent reads — a file, a repository issue, a web page — come back through the server's output and are followed as though the user had typed them.
  • Over-broad credentials. The server runs with the employee's own identity, tokens and access, so least privilege at review time says little about reach at run time.
  • Unreviewed servers added locally. An engineer can install one on their own machine in minutes. That is Shadow AI: agents, MCP servers and Skills running inside the organization that security never approved and cannot see.
Do this But watch out for — and how to contain it
Re-assess each server continuously, not once A stale allowlist entry keeps granting trust after the definition changes; bind approval to the current definition so a change triggers re-assessment
Treat everything a server returns as untrusted input Injected instructions arrive inside legitimate responses; constrain what the agent may do after reading, not only what it may connect to
Scope tokens and access to the task Servers inherit the user's credentials; enforce at the action, where credential reads and unapproved destinations become visible
Keep discovery running on every endpoint New servers appear between reviews; pair ongoing discovery with denylisting and per-team policy

Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome — the record a team needs when an approved connection behaves in a way the original review never anticipated.

Who should own MCP approval, and what evidence proves the process is working?

Which functions own which part of the decision?

No single function should own MCP approval end to end. An MCP server — a Model Context Protocol endpoint that an agent connects to and acts through — touches identity, data scope, and regulatory exposure at once, so the decision splits across owners, and the artifacts each one produces are what make it defensible later.

  • Security owns the risk judgment: the server's publisher, its authentication model, its tool permissions, and whether it is acceptable on an employee machine.
  • IT and MDM administrators own enforcement: turning the decision into an allowlist or denylist entry that applies on the endpoints where agents actually execute.
  • Data governance owns scope: which systems, repositories, and data classes the server's tools may reach.
  • Legal and compliance own external obligations: contractual terms, data residency, and the regulatory regimes the deployment falls under.

What steps produce the audit record?

  1. Register the request with the server name, publisher, transport, tool inventory, and the user identity it will run under.
  2. Attach the risk assessment and the named reviewer from each owning function, with the date of review.
  3. Publish the outcome as an enforceable policy scoped to a team or risk profile, not as a wiki page that nothing reads.
  4. Retain run-level traces so an approved connection can be reconstructed afterward.
  5. Re-review on a defined cadence and on any change to the server's tool set.

By its own account, Backslash Security automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, so these records reach an auditor without a manual collection exercise. Refresh the inventory behind each approval on the same cadence as re-review, and make that inventory cover the whole employee population, since agents, Skills and connectors install under individual user identities on any managed machine.

Frequently Asked Questions

What are you actually vetting when you review an MCP server?

Vetting an MCP server means examining the connection it opens, not just the publisher behind it. MCP, the Model Context Protocol, is the protocol agents use to reach external tools and data sources, and each MCP server is a live path an agent can act through. A workable review covers:

  • Reach — which file paths, repositories, internal APIs, SaaS tenants or cloud credentials the server can touch.
  • Identity — whether it runs under the employee's own account and permissions rather than a scoped service identity.
  • Mutability — whether tool definitions, prompts or endpoints can change after approval without a new review.
  • Install and update path — how it arrived on the machine, and whether updates are pinned or pulled on every run.
  • Destination — where requests terminate, and whether any of that traffic leaves approved infrastructure.

A server that reads a public documentation site and one holding a write-capable token to your Git provider are not comparable risks, even when both end up on the same approved list.

How can you check an MCP server's reputation before approving it?

Public reputation data on an MCP server is the fastest first filter, and it narrows the queue before anyone spends engineering hours on a deep review. As of 22 September 2026, Backslash Security reports that its public MCP Server Security Hub held 81,021 MCP servers, each scored for risk, which gives reviewers a starting score for many of the servers employees install. Reputation alone does not close the review, though: a well-regarded server configured with an over-permissive token, or installed alongside a rules file that quietly redirects its behavior, is still a local problem that no public score can see.

Why doesn't a one-time approval hold?

A one-time approval of an MCP server goes stale because the thing you approved keeps changing — new tools appear in the server, the agent calling it updates, and employees install adjacent components nobody reviewed. This is where Shadow AI shows up: agents, models, MCP servers and Skills running inside the organization that security never approved and cannot see, usually installed by an employee on their own endpoint, sometimes under a personal account. Backslash Security addresses this with continuous discovery of the agentic layer on endpoints plus allowlisting of approved components and denylisting of risky ones, with custom policies per team and risk profile enforced on an ongoing basis rather than at intake.

How do Agent Skills change the picture?

Agent Skills belong in the same review as MCP servers because they extend what an agent can do using the same user permissions. A Skill is packaged instructions and scripts — in practice often just a markdown file on a developer's machine — that can tell an agent to read a file, call an API or run a shell command. Backslash Security operates a free Skills Security Scanner that scans AI agent Skills for security risks, which gives reviewers a way to treat Skills as reviewable artifacts rather than invisible configuration sitting beside the approved tooling.

What do you do about components nobody submitted for review?

Unreviewed components are common, because installing an MCP server or a Skill takes a single command and no administrative approval. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, and blocks risky agent actions inline before execution — unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. For the record afterward, Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC.

Where can a team start without deploying software?

Teams can begin MCP server security work without installing anything permanent. Backslash Security offers a free AI Endpoint Exposure Assessment — the Agentic AI Exposure Assessment (Scout) — which is agentless, read-only and retains no data, distributed through existing mobile device management such as Intune or Jamf. Backslash Security also publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to make the switch?

See why teams choose Backslash Security.

Book a Demo