At a glance
- Before approving a new MCP server, establish ownership, the credentials it inherits, the tools it exposes, data destinations, and post-approval traceability.
- Each MCP server is a live connection an agent can act through, usually carrying the employee's own identity and access.
- Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.
- Backslash Security's MCP Server Security Hub has scanned and scored more than 80,000 publicly available MCP servers.
- Approval is a point-in-time judgment; allowlisting, custom policy and inline blocking before execution keep that judgment valid afterward.
Backslash Security
Published:
Before you approve a new MCP server, work through a short, specific set of questions: who publishes and maintains it, what credentials and permission scopes it will inherit from the person running it, which tools it exposes and what each of those tools is allowed to do, where the data it touches ends up, and what record you will have afterward if something goes wrong. MCP — the Model Context Protocol — is the protocol agents use to connect to external tools and data sources, so every server you approve is a standing connection an agent can act through, typically on an employee's own machine and under that employee's identity. That last detail is what makes the approval decision different from approving a SaaS application: the server does not run in a controlled tenant, it runs on an endpoint with the access its user already holds.
Most approval reviews stall on the same problem, which is that there is no reliable inventory of what is already installed to compare the new request against. Backslash Security addresses that directly by discovering every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, assessing the risk posture of each with an agentless approach, and then blocking risky agent actions inline before execution — including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. The company was founded by Shahar Man, CEO and Co-founder, and Yossi Pik, CTO and Co-founder, to secure the interconnected layer of agents, MCP servers, Skills, plugins, connectors and hooks that employees now run on enterprise machines. As of 2026, that layer has grown faster than the review processes built to govern it, and the questions below are written for teams approving these connections one request at a time.
What does approving a new MCP server actually grant on an employee's machine?
Approving a new MCP server grants an AI agent a standing connection to real systems, operating with the employee's own permissions on their machine. MCP — the Model Context Protocol — is the protocol agents use to reach external tools and data sources: file systems, Git hosts, ticketing systems, internal APIs, cloud accounts. The approval covers everything that channel can reach, for as long as it stays configured.
A server may run as a local process or as a remote endpoint. It usually authenticates with credentials already on the machine — Git configuration, cloud key, package registry token — rather than a separate service identity issued by IT. Backslash Security assesses the risk posture of these endpoint components agentlessly, collecting information without leaving software permanently installed.
Which attributes should a reviewer record before granting approval?
| Attribute | Typical values | Why it decides the answer |
|---|---|---|
| Execution identity | The employee's OS account; a corporate or personal login | Actions land in logs as the human, so audit trails name a person who may not have initiated them |
| Transport and hosting | Local process; remote network endpoint | Determines whether prompts and file contents leave the device, and to whose infrastructure |
| Tool surface | Read file, write file, shell command, API call, database query | Defines the blast radius; shell and write tools behave nothing like read-only lookups |
| Credential scope | Inherited endpoint tokens; dedicated scoped token | Inherited tokens usually exceed what the task requires |
| Update channel | Pinned version; automatic latest | An auto-updating server can gain tools after approval without new review |
| Provenance | Publisher, source repository, maintenance activity | Establishes whether anyone is accountable for what the server ships |
Which questions should you ask about an MCP server's origin, maintainer, and install path?
This depends on what you mean by "where it came from" — the questions worth asking about a new server split into three separate checks. Model Context Protocol (MCP) is the protocol agents use to reach external tools and data sources, and each MCP server is a live connection an agent can act through, so "origin" can mean the published source, the people behind it, or the artifact that actually lands on the machine.
Where is the source of record?
- Which repository or registry publishes it, and is that the project the documentation points to?
- Is this an original project or a fork, and how far has the fork diverged?
- Does the published source correspond to the package that gets built and shipped?
Who maintains it?
- Is the maintainer an organization with a named security contact, or a single unverified account?
- Are releases signed, and is there a visible history of patching reported issues?
- What happens to the server if the maintainer stops responding?
How does it reach the endpoint?
- What is the install path — a package runner such as npx or pip, an IDE extension marketplace, a container, or a script pulled over the network?
- Is the version pinned, or will the next agent run fetch whatever is current?
- Does the server pull extra code or configuration at runtime, after your review has finished?
Useful trust signals are ones a reviewer can inspect directly: signed releases, a published advisory history, and open assessment tooling. Backslash Security publishes Claw-Hunter, an open-source tool for discovering and assessing OpenClaw risks. Record the exact version, package hash, and install path you approved, so the server running later can be compared against the one reviewed.
What should you ask about the tools, permissions, identity, and data an MCP server touches?
Before approving a single MCP server for an employee's endpoint, ask about the tools it exposes, the credentials it uses, the identity it acts under, and the destinations it can reach. This section narrows to that one decision — one server, one approval — rather than the wider governance program.
Ask about the tools first. An MCP server is a Model Context Protocol connection through which an agent reaches external tools and data, and it is approved as one unit, but the individual tools inside it are what actually execute. Each tool's name and description are read by the agent as instructions. Backslash Security supports this review with allowlisting of approved components and denylisting of risky ones, enforced through custom policies tuned to different teams and risk profiles.
| Ask this before approval | Risk if you skip it, and how to contain it |
|---|---|
| Which tools does this server expose, and what does each do? | A server approved once can ship new or re-described tools later. Pin the approved version and re-assess whenever it changes. |
| Which credentials does it use, and whose identity does it act under? | Endpoint servers commonly run under the employee's own token, making their actions indistinguishable from the employee's. Record the identity in your inventory and scope credentials narrowly. |
| What outbound destinations can it reach? | Silent data egress to an unapproved location. Define approved destinations in policy and deny the rest. |
| What externally authored content does it read? | Prompt injection — hidden instructions planted in a file, issue, page or repository the agent reads on its own and follows as though the user had typed them. Treat that content as untrusted and limit which write-capable tools it can reach. |
| Which actions can it take without asking the user? | File writes, shell commands, repository pushes and credential reads happen at machine speed, so after-the-fact alerting arrives too late. Require enforcement that acts before execution. |
How do you decide between approving, approving with constraints, or declining an MCP server?
Deciding between approving, approving with constraints, and declining an MCP server starts with criteria fixed before any individual server reaches review. An MCP server is a connector built on the Model Context Protocol, the standard an agent uses to reach external tools and data, so each one is a live path an agent can act through under the employee's own identity.
Four criteria do most of the work:
- Credential scope — which tokens, keys, or sessions the server can reach. Decisive when the endpoint is a developer workstation holding Git and cloud credentials.
- Action surface — whether the exposed MCP tools only read, or can also write, execute shell commands, or push to production. Decisive when the agent runs unattended.
- Provenance and maintenance — who publishes the server, how it is updated, and whether the installed build matches what was reviewed. Decisive for anything pulled from a public registry.
- Data destinations — where responses and context travel once the agent calls out. Decisive under EU AI Act, NIS2, DORA, or SOC 2 obligations.
| Outcome | What the criteria look like | What it requires operationally |
|---|---|---|
| Approve | Read-only action surface, no standing credential access, known publisher, destinations inside the corporate boundary | Add to the approved allowlist; keep discovery running so version changes are seen |
| Approve with constraints | Useful capability paired with a write or execute path, or credential reach that cannot be narrowed at install time | Scope by policy to specific teams; retain trace data for review |
| Decline | Unverifiable publisher, unconstrained command execution, or egress to destinations you cannot account for | Denylist the component and alert on reinstallation attempts |
Backslash Security supports these outcomes by assessing endpoint agentic components with an agentless approach and letting teams build custom allowlist and denylist policies per team and risk profile, then enforcing them continuously.
What has to happen after approval, when the MCP server is running every day?
After approval, the work shifts from evaluation to continuous control, because an MCP server — the Model Context Protocol connection an agent acts through to reach external tools and data — rarely stays the thing that was reviewed. Versions ship, new tools appear under the same server name, requested scopes widen, and copies spread to teams and endpoints the original request never covered, often under personal accounts rather than corporate identities.
Approval records freeze a server at review time while endpoints execute whatever it has since become.
This stage belongs to the team already operating approved servers day to day, not the team still deciding. The practical sequence:
- Keep discovery running, not scheduled. Backslash Security maintains a live view of the agentic components present on employee machines, so a second install or a newly exposed tool definition surfaces without anyone filing a ticket.
- Re-assess on change. Treat a version bump, an added tool, or a widened credential scope as a reassessment trigger in its own right, rather than waiting for a quarterly review window.
- Express the decision as enforceable policy. Allowlist what you cleared, denylist what you rejected, and write policies per team so groups with different risk profiles are governed differently.
- Extend control to runtime behavior. A catalog entry records which server is permitted; it says nothing about what an approved server is driven to do mid-run, so controls have to reach the moment an agent attempts an action on the endpoint.
- Preserve a reconstructable record. Per Backslash Security, the platform traces the full path of an agent run from prompt to agent to tool call to outcome — the record an investigation needs afterward.
Each step applies across all employees and all endpoints, including machines that never filed the original request.
Frequently Asked Questions
What should I ask before approving a new MCP server?
Before approving a new MCP server — a connector built on the Model Context Protocol, the standard agents use to reach external tools and data — work through a short, consistent set of questions:
- Whose identity does it act under? Most endpoint MCP servers run with the employee's own credentials and access, not a service account.
- What tools and scopes does it expose? Each tool is an action an agent can take without a human in the loop.
- Where does data go? Name the destinations it may contact, and treat everything else as unapproved.
- Who publishes and updates it? An unattended auto-update changes the thing you approved.
- What record survives the run? Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which is what an investigation later needs.
Why doesn't my existing endpoint stack answer these questions?
Endpoint detection and response (EDR) was built to catch malicious processes, files and known-bad behavior on a machine, and mobile device management (MDM) controls devices and installed applications. Neither layer reads the instructions an agent was given, the MCP tools it can call, or the Agent Skills — packaged instructions and scripts that extend what an agent can do, often just a markdown file on the machine — that it loads at runtime. The process looks legitimate because it is legitimate. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, which is the inventory this review depends on.
How do I find the MCP servers nobody asked permission for?
That is Shadow AI: AI tools, agents, MCP servers and Skills running inside the organization that security has not approved and cannot see, usually installed by employees on their own endpoints and often under personal accounts such as a private Gmail login on a corporate machine. You cannot approve what you have not enumerated, so discovery comes before policy. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless — meaning it collects information without leaving software permanently installed on the endpoint — read-only, and retains no data, giving reviewers a starting inventory rather than a guess.
Is a reputable MCP server safe once approved?
Approval settles provenance, not behavior. An approved agent can still become a rogue agent — reaching for credentials, escalating privilege, or sending data to a destination nobody sanctioned — typically after it is manipulated. Prompt injection is the common route: hidden instructions placed in content the agent reads, such as a file, repository, issue or web page, which it follows as though you had typed them. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
What evidence should an approval decision leave behind?
Auditors and incident responders ask the same thing after the fact: what was approved, by whom, and what did it actually do. Record the requester, the scopes granted, the allowlist or denylist entry created, and the policy that governs it — Backslash Security supports allowlisting of approved components, denylisting of risky ones, and custom policies tuned to different teams and risk profiles. According to Backslash Security, the platform automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, so the review trail is produced as a by-product of enforcement rather than reconstructed later from memory.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07