At a glance
- Denylisting refuses known-bad components before load; sandboxing limits blast radius; inline blocking stops a specific agent action in the moment before execution.
- None of the three works without discovery first — you cannot denylist or allowlist components nobody has inventoried on the endpoint.
- Backslash Security research found malicious repository configurations can redirect Claude Code API requests to an attacker-controlled server, silently exposing the user's sign-in token.
- Shadow AI — unapproved agents, MCP servers and Skills installed by employees under personal accounts — is the population these controls are usually aimed at.
- Match the control to the risk profile of each team rather than standardizing one mechanism across every endpoint.
Backslash Security
Published:
Denylisting, sandboxing, and inline blocking are three distinct controls for the AI plugins employees install on their machines, and they act at three different moments. A denylist refuses a named component — a specific MCP server, Skill, or extension — before it ever loads, which works when the risky thing is already known by name. A sandbox confines what a component can reach once it is running, limiting damage without understanding what the agent was asked to do. Inline blocking evaluates the action an agent is about to take and stops it in the moment before execution, which is the only one of the three that reacts to behavior rather than identity.
Some definitions, since the vocabulary is young. MCP is the Model Context Protocol, the standard agents use to connect to external tools and data sources; each MCP server is a live connection an agent can act through. Agent Skills are packaged instructions and scripts that extend what an agent can do — often a single markdown file on a developer's machine — and they execute with that user's own permissions. Shadow AI covers any of this that security never approved and cannot see.
Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, and blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. As of 2026, that combination — inventory first, then policy, then enforcement at the action boundary — is what the control decision below is built around.
What is an AI plugin control, and what is it actually controlling?
This section narrows to one layer: the controls applied to AI components running on employee endpoints, not to models hosted in a cloud service. An AI plugin control is any enforcement mechanism that decides whether a given piece of the agentic AI fabric — the interconnected layer of agents, models, MCP servers, MCP tools, Skills, hooks, rules files, plugins, connectors and context on a machine — may load, run, or act under that employee's identity. Three archetypes dominate: denylist, sandbox, and inline block.
Some vocabulary, since these objects are new to most control frameworks:
- MCP (Model Context Protocol) — the protocol agents use to reach external tools and data; each MCP server is a live connection an agent can act through.
- Agent Skills — packaged instructions and scripts that extend an agent, often just a markdown file on the machine, executing with the user's own permissions.
- Hook — a trigger that fires on an agent action, running something before or after it.
- Rules file (AGENTS.md, CLAUDE.md) — standing instructions an agent reads on every run.
Each control archetype is best described by its attributes:
| Attribute | Denylist | Sandbox | Inline block |
|---|---|---|---|
| Decision point | Install or load time | Execution environment | Moment of the action |
| Input it needs | Known-bad identifiers | Isolation boundary | The pending action and its run context |
| Granularity | Whole component | Whole session | Individual action |
| Fails toward | Unknown items allowed | Broad restriction | Action-level judgment |
Why the attributes matter: a denylist cannot judge a component it has never seen, a sandbox constrains the environment the agent runs inside, and action-level enforcement depends on visibility into what the agent is now attempting and the run that produced it. Backslash Security provides allowlisting and denylisting of components and inline blocking of risky actions on the endpoint, where agents, MCP servers and Skills execute with the employee's own permissions and access.
How do denylisting, sandboxing, and inline blocking differ in practice?
Denylisting, sandboxing, and inline enforcement intervene at different moments in an agent's life cycle, which is why they produce different coverage and different audit evidence. Denylisting names components that are not permitted to run — a specific MCP server, model, or Agent Skill. Sandboxing isolates execution inside a restricted environment with narrowed file, network, and credential access. Inline blocking intercepts an individual agent action as it is attempted and stops it before it executes.
Five criteria separate them in practice:
- Enforcement point — whether the control sits on the component inventory, at the operating-system boundary, or on the agent's action path. Decisive when agents run under an employee's own identity and inherit that person's access.
- Timing relative to execution — pre-install, during process execution, or at the tool call. Matters because a credential read or an outbound transfer cannot be reversed once completed.
- Coverage of unknown items — whether a never-before-seen Skill, hook, or rules file is handled by default.
- User friction — how often legitimate engineering work is interrupted.
- Evidence produced — what an investigator or auditor can reconstruct afterward.
| Criterion | Denylisting | Sandboxing | Inline blocking |
|---|---|---|---|
| Enforcement point | Component inventory and policy | OS / container boundary | Agent action path (tool call, command, destination) |
| Timing | Before install or launch | During execution, inside a confined scope | At the instant of the action, pre-execution |
| Unknown items | Not covered until catalogued | Contained, but behavior still unobserved | Evaluated on behavior, not prior knowledge |
| User friction | Low, but stale lists block approved work | High where real repositories, credentials, and tooling are needed | Targeted — only the violating action stops |
| Evidence | Policy decision records | Container logs, limited run context | Action-level record of what was attempted and refused |
Denylisting fits catalogued, clearly disallowed components; sandboxing fits workloads that tolerate isolation; action-level enforcement fits approved agents whose individual steps need constraining. On the evidence criterion, Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome.
Why does a denylist fall behind the pace of new plugins and MCP servers?
A denylist blocks only what someone has already named, which is why it falls behind as new plugins, Skills and MCP servers — Model Context Protocol servers, the connections an agent acts through — appear on employee machines week after week. An allowlist inverts the logic but inherits the same dependency: both are string-matching controls applied to an inventory nobody fully holds. This means the list can only ever cover what discovery has already surfaced.
Three structural limits show up consistently:
- The discovery gap. Components installed by employees under their own accounts — shadow AI — sit outside the list by definition, because a control can only deny what was observed in the first place.
- Naming and versioning drift. An Agent Skill is often a packaged set of instructions in a markdown file; forked, renamed or bumped a version, the same capability no longer matches the entry that blocked it.
- The identity question. Lists record components, not people. Without the endpoint and account that introduced a plugin or MCP server, there is no record to reconstruct afterward.
Backslash Security's MCP Server Security Hub, a public and continuously updated risk database, held 81,021 MCP servers when read on 22 September 2026, each scored for risk, giving policy a risk posture to anchor to when names and versions move.
| Do this | But watch out for — and how to offset it |
|---|---|
| Denylist components already judged risky | Coverage stops at what has been named; pair it with continuous discovery so new installs enter the record as they arrive |
| Allowlist approved agents and servers | Forks and version bumps slip the match; anchor approval to assessed risk posture, reassessed continuously |
| Treat install events as a control point | Personal-account installs leave no owner on the finding; bind each component to the endpoint and identity behind it |
When does sandboxing an AI plugin help, and where does it stop helping?
Sandboxing an AI plugin helps most when the component's work is genuinely self-contained, and it stops helping the moment the agent needs real access to do its job. This depends on what you mean by sandboxing. In the strict sense it is OS-level isolation — a container or restricted process with a narrowed filesystem view, no privileged process spawning, and controlled network egress. In the looser sense it means running an extension in a throwaway workspace. Only the strict form constrains anything an attacker cares about, and even that form limits where code can reach while leaving the instruction path untouched.
| Do this | But watch out for this |
|---|---|
| Run unreviewed plugins and Agent Skills — packaged instructions and scripts that extend an agent, often just a markdown file — in a restricted execution environment | Isolation limits where a shell command lands, not whether the agent reads instructions telling it to run one; pair the sandbox with an assessment of what the component may invoke |
| Scope the filesystem view to a single working directory | Local repositories, Git configuration and token files usually sit inside that directory; mount read-only where possible and keep secrets out of the tree |
| Restrict outbound traffic to approved destinations | An MCP server — the protocol connection an agent acts through — is an approved channel that can still carry data outward; govern connectors by allowlist, not by port |
Isolation breaks down for the work employees actually want. Agents operate through corporate SaaS connectors brokered by identity providers such as Okta or Entra ID, under the employee's own account, against local repositories that employee is entitled to read. A sandbox permitting that access cannot separate authorized use from an agent that has drifted past its assigned task, because both arrive with the same credentials. Backslash Security works at the action layer, evaluating what an agent is attempting with the access it already holds.
What does blocking a risky action inline before execution give you that the other two cannot?
When an agent is already mid-run, blocking a risky action inline is the only control that judges the specific tool call in front of it, instead of a category decided weeks earlier. Denylisting decides at the component level — which agents, MCP servers and Skills may exist on the machine. Sandboxing decides at the environment level — what a process can reach. Action-level enforcement decides at the level of the individual call an agent is about to make, with the run's context attached.
The attributes that define this control layer:
- Decision point — at the moment of execution, before the call lands. Backslash Security acts before the violation executes, which matters because an approved component can still attempt an unapproved action.
- Unit evaluated — one tool call together with the prompt, agent identity and target that produced it, not a package name or a directory boundary.
- Outcomes — allow or stop. The safe path proceeds untouched, so engineers do not experience governance as friction.
- Scope — all employees and all endpoints, with developer workstations a prominent use case inside that scope.
- Residue — every decision leaves a record of what was attempted, which is what makes an incident reconstructable afterward.
Evaluation at the moment of execution carries information no earlier checkpoint holds: the same tool call that is routine in one run is credential exfiltration in another, and only the surrounding run distinguishes them.
Practically, the five steps run in order — discover what is installed, assess its posture, govern it with allowlists and denylists, block the risky call inline, then reconstruct the path from prompt to outcome.
Frequently Asked Questions
What is the difference between denylisting an AI plugin and blocking an agent action inline?
Denylisting an AI plugin and blocking an agent action inline are decisions taken at different moments. A denylist is a component-level judgment made before anything runs: a named plugin, Model Context Protocol server, or Agent Skill — a packaged instruction set that executes with the user's own permissions — is marked unapproved and kept off the endpoint. An inline block is an action-level judgment taken while the agent is working. Backslash Security supports allowlisting of approved components and denylisting of risky ones, with custom policies for different teams and risk profiles, and blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
Why is a denylist not sufficient on its own?
A denylist governs which components exist on the machine; it carries no judgment about what an approved component does once it runs. Backslash Security research discovered that malicious repository configurations can redirect Claude Code API requests to an attacker-controlled server, silently exposing the user's sign-in token — allowing attackers to run AI workloads, access profile information, upload files, and initiate sessions under the victim's identity. The agent in that scenario is one an allowlist would happily permit. Configuration and context, including rules files such as AGENTS.md or CLAUDE.md that hand an agent standing instructions on every run, change behavior without changing the inventory.
How does sandboxing compare with inline prevention for AI plugins?
Sandboxing and inline prevention answer different questions about an AI plugin. A sandbox is a general isolation technique that limits which files, processes, and network destinations a workload can reach. It constrains blast radius, and it has no visibility into the task the agent was asked to perform or the tool calls it chose. Inline prevention evaluates the pending action itself and stops it before it executes. Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills.
Which components should be inventoried before any control is chosen?
Every control decision depends on an inventory, and in agentic AI endpoint security that inventory is wider than a plugin list. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin, and connector running on an endpoint. That discovery also surfaces shadow AI — unapproved agents, models, MCP servers, and Skills that security cannot see, frequently running under an employee's personal account on a corporate machine. Backslash Security offers a free AI Endpoint Exposure Assessment, called Scout, which is agentless, read-only, and retains no data. Philip Walsh, Head of Security Engineering at Happy Returns, said: "Backslash gave us a live picture of the AI our engineers were already running, and a way to govern it without slowing anyone down."
What happens when an approved agent starts acting outside its assigned task?
An approved agent acting outside its task is a rogue agent: one that reaches for credentials, escalates privilege, or sends data to an unapproved destination, often because it was manipulated by injected instructions. Neither a denylist nor an isolation boundary addresses this, because the component is sanctioned and the activity is its own. Backslash Security's real-time protection acts on the endpoint before the action executes, so a sanctioned agent reaching for credentials or an unapproved destination is stopped at that call rather than discovered afterward.
How do teams evidence these plugin controls to auditors?
Evidencing plugin controls requires a record of what ran, under which policy, and what was stopped. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, which gives incident responders a reconstruction rather than a process-level log fragment. Backslash Security also automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, so the allowlist, denylist, and inline prevention decisions applied on each endpoint can be presented as compliance artifacts.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07