Comparison

How to Evaluate an MCP Server's Risk Before Employees Install It

At a glance

  • Evaluate an MCP server on publisher provenance, exposed tools, inherited credentials, exposure to untrusted content, and whether you can still see it post-install.
  • Each MCP server runs with the installing employee's own identity and permissions, so an endpoint-level approval decision carries enterprise-level consequences.
  • Backslash Security was founded to secure the agentic AI fabric: agents, MCP servers, Skills, plugins, connectors and hooks running on enterprise endpoints.
  • Backslash Security discovers and scores those components, allowlists approved ones, and blocks risky agent actions inline before they execute.

Backslash Security

Published:

Evaluating an MCP server's risk before employees install it comes down to a short set of concrete checks: who publishes and maintains the server, which tools it exposes and what each of those tools is permitted to do, which credentials, tokens and file paths it inherits from the person running it, whether its inputs can be influenced by content your organization does not control, and whether security will still be able to see it once it is installed. MCP, the Model Context Protocol, is the protocol AI agents use to reach external tools and data sources, which means every MCP server is a live connection an agent can act through — and it runs with the installing employee's own identity and permissions, not a scoped service account. Answering those questions in advance is what separates a reviewed connection from an unreviewed one.

In practice the decision happens on an endpoint, in a configuration file, when an engineer adds a server to Claude Code, Cursor or another agent client without opening a ticket. The same machine usually holds Agent Skills — reusable capabilities an agent can invoke, often just a markdown file sitting in a folder — along with hooks, rules files and plugins that shape what the agent does next. Backslash Security was built for that layer: it continuously discovers the agents, models, MCP servers, MCP tools, Skills, hooks, rules files, plugins and connectors running on employee machines, assesses their risk posture using an agentless approach, and lets security teams allowlist what is approved and denylist what is not. As a reference point for the servers employees are most likely to reach for, Backslash Security reports that its MCP Server Security Hub has scanned and scored more than 80,000 publicly available MCP servers, a figure read on 22 September 2026. And when an approved component behaves in a way nobody asked for, Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome.

What makes an MCP server risky before anyone installs it?

Narrowing to the pre-install decision alone: what makes an MCP server risky is largely knowable before the install, from properties you can read off the server itself. MCP — the Model Context Protocol — is how an AI agent connects to external tools and data sources, so each MCP server is a live connection an agent can act through. On an employee machine, it usually runs as a local process started by a configuration file, under that employee's own identity, tokens, and file permissions. Nothing elevates; the server simply inherits the user.

That inheritance is why a handful of attributes determine the AI risk before anything executes.

Execution location. Values: local process on the host, or remote endpoint. A local server runs code on the machine with the user's access; a remote one moves data off it. Each implies a different control point.

Credential scope. Values: none, scoped token, or ambient access to environment variables, cloud credential files, and Git configuration. Ambient access is the difference between a tool that reads a calendar and one that can reach production.

Tool surface. Values: read-only queries, file writes, shell command execution, outbound network calls. Shell and egress tools are what turn a drifting agent into a real incident.

Untrusted-content exposure. Values: trusted inputs only, or content the server fetches from repositories, issues, web pages, and documents. The latter is the delivery route for prompt injection — hidden instructions inside content the agent reads, which it follows as though the user had typed them.

Provenance and update behavior. Values: pinned version from a known publisher, or an auto-updating package from an unverified source. An unpinned server can change behavior after approval.

Backslash Security assesses the risk posture of endpoint agentic AI components using an agentless approach, meaning the assessment collects this information without leaving software permanently installed on the machine.

Which signals should a pre-install MCP server risk review collect?

A pre-install review of an MCP server collects signals before the server runs on a work machine. The scope is narrow: one Model Context Protocol server—the connection an AI agent uses to reach external tools and data—assessed before first execution.

Signal What the reviewer collects Why it matters
Publisher identity Repository owner, maintainer history, package namespace, signing status An unverified or recently created publisher is a cheap entry point for supply chain attacks
Tool manifest The full list of tools the server exposes, with each tool's description text Tool descriptions are read by the agent as instructions, so the manifest is executable surface, not documentation
Declared scopes Filesystem paths, API permissions, shell access, read versus write The server inherits the employee's identity and access; scope is the blast radius
Network destinations Hostnames and endpoints the server contacts, including telemetry Shows where enterprise context can leave the endpoint
Credential handling Where tokens are stored, whether secrets are read from environment variables or config files Determines exposure of cloud keys, package registry tokens and Git configuration
Update mechanism Pinned version, floating tag, or self-updating install A floating dependency means today's approval governs code nobody has reviewed
Prompt-injection surface Whether the server returns third-party content—issues, pages, documents—into the agent's context Prompt injection is hidden instructions in content an agent reads, which it follows as though the user typed them

Tool manifests and update mechanisms can change without reinstall, so recorded decisions age between reviews unless re-checked. Backslash Security assesses endpoint agentic AI components using an agentless approach—collecting information without permanently installing software—so these attributes are re-evaluated continuously rather than captured once. Privately built and internal servers have no public reputation, so every signal must be gathered directly from code and configuration.

How do the control layers that can evaluate MCP servers compare?

Four control layers evaluate an MCP server—the connection through which an agent reaches an external tool or data source—before employee installation, each answering a different question:

  • Pre-install visibility: does the layer reveal a server exists and its capabilities before agent connection? Decisive for employee-driven adoption without tickets.
  • Enforcement point: where decisions apply—human judgment, list, network path, or agent host. Decisive when risky calls must be stopped, not logged.
  • Durability under change: how the layer withstands server updates, forks, or replacements.
  • Forensic record: whether the layer leaves incident-reconstruction evidence.
Layer Pre-install visibility Enforcement point Durability under change Forensic record
Manual review Strong for submitted servers; none otherwise Human judgment and policy Re-review needed per version Review notes only
Registry or allowlist curation Strong for curated entries The list, applied at request time Drifts between curation cycles Approval history
Network-side proxying Inferred from traffic once connections begin Network path; suits environments where endpoint deployment is difficult Holds for routed traffic Connection-level telemetry
Endpoint-side discovery with inline action control Direct, from agent's machine The host, at agentic layer Continuous rediscovery as components change Host-side activity trail

Manual review suits small, slow-moving server sets. Curated registries suit organizations funding standing review functions. Network-side proxying suits estates with constrained endpoint deployment. Endpoint-side discovery suits organizations where employees install components under their own identities—the case Backslash Security addresses, pairing agentless discovery (collection leaving nothing permanently installed) with on-host guardrails acting on agent actions before execution.

Why do existing endpoint and identity controls miss employee-installed MCP servers?

When an engineer installs an MCP server on a corporate endpoint under their own identity, software procurement controls don't register the event. MCP, the Model Context Protocol, connects agents to external tools and data sources—typically a few configuration file lines rather than a signed installer. Nothing was purchased, packaged, or triggered an admin prompt.

The gaps are structural:

  • MDM inventories applications, not agent configuration. Intune or Jamf detect Claude Code or Cursor; the MCP servers, Agent Skills, hooks and rules files registered inside sit below the inventory boundary.
  • Identity providers authenticate the human. Okta or Entra ID issue employee sessions, and agents act with that employee's access, so authorization logs show approved people doing approved things.
  • EDR resolves processes. Endpoint detection catches malicious files and known-bad behavior; it records legitimate interpreters executing while the prompt, Skill or server causing the call stays outside its view.
Do this Watch out for How to mitigate it
Pull an MDM report of installed AI clients It surfaces the editor while connected servers remain invisible Enumerate the configuration layer on each machine, component by component
Apply application allowlisting A tool server often runs as a script invoked by an already-approved binary Allowlist and denylist agentic components themselves, with policies per team
Inspect agent traffic at the network layer Local models and loopback tool calls never cross the perimeter Pair network inspection with host-side visibility
Review identity logs after an incident Agent actions are indistinguishable from the user's own Keep tracing data linking a prompt to the tool call it produced

Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.

What does a repeatable MCP server risk scoring workflow look like?

A repeatable MCP server review turns an ad hoc install request into a documented sequence with a named owner and an exit condition at each stage. MCP, the Model Context Protocol, is how an agent connects to external tools and data, so every server added is a new path the agent can act through. This section addresses the consideration stage—when an organization has accepted employees are already running agents and is deciding how to govern them.

If a connector must be judged before installation, the organization first must know what is already installed. Discovery comes before policy, not after.

The stages

  1. Intake. Capture the requester, business purpose, publisher and origin of the server, plus systems it will reach.
  2. Discovery. Baseline what endpoints already run—agents, MCP tools, Agent Skills, hooks, rules files such as AGENTS.md or CLAUDE.md, plugins and connectors. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data.
  3. Static assessment. Read the manifest and declared tools, credentials and scopes requested, authentication and transport model, and publisher's provenance.
  4. Tiering. Sort candidates by reach: read-only public data, internal systems, or credentialed write access to source control and production.
  5. Conditional approval. Approve against a policy scoped to the requesting team's risk profile, with explicit conditions on which tools may be invoked.
  6. Enforcement. Apply guardrails on the host, at the agentic layer where the agent executes, so the approval decision is enforced when the agent acts.
  7. Reconstruction. Retain audit and tracing records of agentic activity so an investigator can establish afterward what the agent did and why.

Who should own MCP approval, and what evidence proves the review happened?

No single team owns MCP server approval end to end: security sets the risk bar, the IT or MDM administrator controls how components reach the machine, and the AI governance lead reports against the resulting inventory. What makes that split workable is naming one accountable owner for the decision and treating the other two parties as suppliers of evidence to it.

What steps make ownership and policy concrete?

  1. Name the accountable owner in writing. One role signs off on each MCP server, Skill and rules file — a file carrying standing instructions an agent reads on every run — even when discovery and enforcement sit with other teams.
  2. Write policy against components, not vendors. State which agents, servers, Skills, hooks and connectors are allowlisted, denylisted, or need review, so the rule survives the next tool your engineers adopt.
  3. Publish the decision where employees already are. A request path in Jira or ServiceNow, with the current allowlist visible, turns policy from a document into a usable route.
  4. Record the evidence with the decision. Server identity, version reviewed, permissions granted, approver, and assessment date.
  5. Set a re-review trigger, not just a calendar. Reassess on version change, new tool exposure, or ownership transfer of the upstream repository.

What evidence demonstrates the review is current?

Audit value comes from traceable records rather than attestations. According to Backslash Security, the platform automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, giving governance leads something to report against without manual collection.

Approval records age at the speed of the components they describe, so a decision captured without its version and read date stops working as evidence the moment the server ships an update. Through 2026, each model release and protocol change across the connector and Skills surface restarts that clock.

Frequently Asked Questions

What should we check before approving an MCP server for employee endpoints?

An MCP server — a connector built on the Model Context Protocol, the standard agents use to reach external tools and data — should be reviewed on provenance, permissions, and reach. Practical checks: who publishes and maintains it, which tools it exposes, what credentials or tokens it requires, which destinations it can write to, and whether it can execute shell commands. Backslash Security runs an MCP Server Security Hub that, as read on 22 September 2026, has scanned and scored more than 80,000 publicly available MCP servers, which gives reviewers a scored starting point instead of a blank page.

Why isn't a one-time review of an MCP server enough?

Because the package you approved is not the only thing that determines what an agent does with it. Prompt injection — hidden instructions placed in a file, repository, issue or web page that an agent reads and follows as if a user had typed them — can redirect an approved connector toward actions nobody requested. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Connectors also update, change scopes, and acquire new tools after approval, so posture needs continuous reassessment rather than a single gate.

How is an MCP server different from an Agent Skill, a rules file or a hook?

They are separate components of the same layer. An MCP server is a connection an agent acts through. Agent Skills are packaged instructions and scripts — often a markdown file on the user's machine — that execute with the user's own permissions. A rules file such as AGENTS.md or CLAUDE.md carries standing instructions the agent reads on every run. A hook is a trigger that fires before or after an agent action. Backslash Security was founded to secure the agentic AI fabric: the mesh of agents, MCP servers, Skills, plugins, connectors and hooks running on enterprise endpoints, because none of these is meaningfully assessed in isolation.

Can MDM or our existing endpoint tooling show which MCP servers employees installed?

Usually not at this level of detail. MDM platforms such as Intune and Jamf manage applications, profiles and device configuration; an MCP server is frequently a configuration entry plus a package a user installs under their own identity, below that management boundary. Endpoint detection and response tooling is built to catch malicious processes, files and known-bad behavior, so it resolves the process that ran — the agent, Skill or connector that caused the action sits a layer above. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which is what produces a reportable inventory.

How can we assess our exposure without deploying software to every laptop first?

Start agentless — meaning information is collected without leaving software permanently installed on the machine. Backslash Security offers a free AI Endpoint Exposure Assessment, called Scout, that is agentless, read-only, and retains no data; it is distributed through your existing MDM, runs, reports and finishes. That gives discovery and risk posture for the agentic components already present. Enforcement is a separate matter: blocking an action on the host requires on-host capability, so treat the assessment as the inventory step that tells you whether deeper controls are warranted.

What happens when an employee is already running an unapproved MCP server?

This is Shadow AI — AI tools, agents, MCP servers and Skills running inside the organization that security has not approved and cannot see, often installed under a personal account. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations, and traces the full path of an agent run from prompt to agent to tool call to outcome, so an investigation has a record to work from. Backslash Security also automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. As Philip Walsh, Head of Security Engineering at Happy Returns, put it: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-08

Ready to make the switch?

See why teams choose Backslash Security.

Book a Demo