Blog

What Makes a Third-Party AI Plugin Risky on Employee Laptops?

At a glance

  • A third-party AI plugin inherits the employee's identity, permissions and access, so it can reach whatever systems, repositories and credentials that person can.
  • Hidden instructions inside content an agent reads can redirect an already-approved component toward credential access or data exfiltration, with no malware involved.
  • Where no inventory of installed agents, MCP servers, Skills, hooks and rules files exists, security owns risk it cannot measure or report.
  • Backslash Security discovers every agent, MCP server, Skill, hook, rules file, plugin and connector on an endpoint and blocks risky actions inline before execution.

Backslash Security

Published:

A third-party AI plugin is risky on an employee laptop because it runs with that employee's own identity, permissions and access — not in a sandbox, and not under a reviewed service account. Whether the component is an extension to an AI coding agent, an MCP server (the Model Context Protocol connection an agent acts through), or an Agent Skill (packaged instructions and scripts that extend what an agent can do, often just a markdown file sitting on the machine), it can read local files, call APIs and run shell commands at the privilege level of the person who installed it. Nobody approved it, no inventory records it, and the security tooling already on the machine was built to catch malicious processes and files rather than to see what an agent was instructed to do.

The second half of the problem is that these components take direction from content they read. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file — a rules file carrying standing instructions an agent reads on every run — could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Nothing in that chain looks like an attack at the process level: an approved agent read a file and followed it. Add Shadow AI, meaning AI tools, agents, MCP servers and Skills that security hasn't approved and can't see, often installed under a personal account on a corporate machine, and the exposure widens past the single plugin anyone was worried about. Backslash Security addresses this by discovering every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint, assessing what is safe to use, and blocking risky actions inline before execution — including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations. Its coverage as of 2026 spans AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity. Latio's 2026 AI Security Market Report found that "where Backslash stood out in our evaluation is in control depth" — beyond scanning for what's malicious, the platform provides in-depth control over endpoint agent settings and what is available to an agent in the first place, with real-time protection.

What is a third-party AI plugin on an employee laptop, and what does it actually touch?

A third-party AI plugin on an employee laptop is any component an AI assistant or agent loads from outside the organization and runs locally, under that employee's own identity and access. The phrase carries two distinct meanings in practice, and they point at different controls.

Application extensions. The familiar sense: an add-on installed into an application the employee already uses — an assistant extension in an IDE, a browser add-on, a marketplace connector linking a desktop AI client to a SaaS account. The host application mediates what the extension can do, and installation usually leaves a visible artifact.

Agent-loaded capabilities. This is the sense used throughout this article, because it is where the local blast radius lives. These are the pieces an agent such as Claude Code, Cursor, Codex or GitHub Copilot picks up at run time:

  • MCP servers and MCP tools — Model Context Protocol is the protocol agents use to reach external tools and data; each server is a live connection the agent can act through.
  • Agent Skills — packaged instructions and scripts, often just a markdown file, that can read files, call APIs or run shell commands with the user's permissions.
  • Hooks — triggers that fire before or after an agent action and run something else.
  • Rules files such as AGENTS.md or CLAUDE.md — standing instructions in a repository or on the machine that the agent reads on every run.
  • Plugins and connectors — the glue binding an agent to local or remote systems.

Backslash Security was founded to secure this interconnected layer on enterprise endpoints, where components invoke one another and none is meaningful in isolation. Once loaded, they inherit the laptop's reach: local source trees, Git configuration, cloud and package registry credentials, signed-in corporate applications, and outbound network destinations.

Why does a plugin that runs under an employee's own identity change the risk picture?

When a third-party AI plugin runs under an employee's own identity, it inherits that person's tokens, sessions and entitlements instead of receiving a scoped service account of its own. Installation happens locally — a browser or editor extension, an MCP server (Model Context Protocol, the protocol agents use to reach external tools and data), or a Skill dropped into a home directory. From there the component can read whatever the signed-in user can read: cached cloud CLI credentials, Git tokens, an active Okta or Entra ID session, mounted drives, internal wikis.

This means the authorization path security actually governs is skipped. Conditional access, app registration review and per-application consent all assume a request arrives as a distinct client. A locally installed component arrives as the human. Nothing is spoofed and nothing looks anomalous, because the activity genuinely is the user's. The practical consequence is that the blast radius equals the employee's full entitlement set — a stated purpose describes intended behavior, while the borrowed identity describes what is reachable.

Do this But watch out for — and how to contain it
Inventory the plugins, Skills, MCP servers and connectors on each endpoint A point-in-time list ages immediately; pair it with continuous discovery so new installs surface as they appear
Approve components on stated function Stated function says nothing about the host identity; review approvals against the permissions the signed-in user already holds
Restrict personal logins on corporate machines Policy alone does not reveal them; Backslash Security detects the use of private account access on enterprise machines

Which specific plugin behaviors should a security team treat as risky?

A security team can evaluate specific plugin behaviors directly, and the patterns below are risky by default. This section covers one sub-case: third-party extensions installed into an AI agent on an employee laptop—Skills, MCP servers, connectors and plugins for tools such as Claude Code, Cursor or GitHub Copilot—rather than the agent itself. Three terms recur. A rules file (AGENTS.md or CLAUDE.md) is a file in a repository or on a machine carrying standing instructions the agent reads on every run. A hook is a trigger that fires before or after an agent action. An MCP server is a connection the agent can act through, using the Model Context Protocol.

Behavior What it looks like on the endpoint Why it matters
Silent or excessive permission grants The extension inherits the employee's shell, file system, cloud and Git access with no separate consent step The blast radius equals the user's own identity, not a sandboxed service account
Behavior change after approval An auto-updating package adds tool calls or network destinations after the version that was reviewed The approval record no longer describes what is running
Injection paths via rules files and hooks Hidden instructions sit in a file, issue or page the agent reads on its own The agent follows them as though the operator had typed them
Unvetted or typosquatted registry entries A near-name match to a popular Skill or server, installed from a public registry Installation is a one-line command with no provenance check
Egress to unknown model endpoints Prompts, file contents or secrets routed to an inference endpoint nobody approved Data leaves through an application-layer path
Unconfirmed chained tool calls One call triggers the next—read, then fetch, then write—with no human gate Actions execute faster than review can intervene

Backslash Security assesses these components on the endpoint and marks which are safe to use.

Why do existing endpoint and network controls miss this layer?

When a third-party AI plugin runs on an employee laptop, existing endpoint and network controls see an approved application doing ordinary work. The agent process is signed and familiar, and its outbound session goes to a permitted model provider. What makes the action risky—which tool the agent invoked, which instructions it read, and what it was attempting—happens inside that session.

Criteria that decide whether a control can govern agentic activity:

  • Unit of observation — whether the layer resolves processes and packets, or individual tool invocations and the context an agent was given. A single approved process can carry many distinct agent actions.
  • Decision point — whether enforcement can occur before an action executes or only after telemetry is written. This is decisive for irreversible actions such as credential access or data egress.
  • Identity context — whether the layer can distinguish corporate identity from a personal login used on a company machine.
  • Evidence retained — whether the record supports reconstructing an agent run later, which is what compliance and incident investigation need.
Control layer What it inspects What stays out of view
Endpoint detection and response Processes, files, known-bad behavior on the machine The instructions an agent read and the tool calls it made inside an approved process
Network and proxy inspection Destinations, traffic patterns, content in transit Why a permitted model endpoint was contacted, and under whose account
Device management (MDM) Installed applications, configuration posture Skills, hooks and rules files added to a user's home directory
Identity and access tooling Logins, tokens, entitlements What an agent does once it inherits the user's own permissions

How can an organization discover, assess and govern AI plugins before they cause harm?

An organization can discover and assess third-party AI plugins before they cause harm by running one repeating sequence across every endpoint, rather than a one-time audit.

What does the sequence look like, step by step?

  1. Inventory the whole layer, continuously. Enumerate agents, models, MCP servers (Model Context Protocol endpoints an agent connects through to reach external tools and data), MCP tools, Agent Skills (packaged instructions and scripts that extend what an agent can do, running with the user's own permissions), hooks (triggers that fire before or after an agent action), rules files, plugins and connectors. Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data.
  2. Assess risk posture per component. Score what is already installed: where it reaches, what permissions it inherits, and whether its origin is known. Unsanctioned AI use — a server nobody approved, or a personal account signed in on a corporate machine — surfaces here as a named list.
  3. Assign policy by team. Allowlist approved components, denylist risky ones, and write separate policies so an engineering group and a finance group are not governed by the same rule set.
  4. Decide at execution time. Backslash Security evaluates an agent's attempted action and stops the risky ones before they run.
  5. Keep a reconstructable record. Retain tracing of agent activity so an investigation can follow what ran, under whose identity, and what it touched.

Endpoint AI exposure is governed at the point of action, not at the point of installation — a component that looked benign in the catalog behaves differently the moment it reads attacker-supplied content. New Skills and servers keep appearing on endpoints as of 2026, which is why the first step runs continuously rather than on an audit calendar.

Frequently Asked Questions

What makes a third-party AI plugin risky on an employee laptop?

A third-party AI plugin — any connector, extension, MCP server or Agent Skill that adds capability to an AI assistant — is risky because it executes with the employee's own identity, permissions and network access, not inside a sandbox. Model Context Protocol (MCP) is the protocol agents use to reach external tools and data, so each MCP server is a live path an agent can act through. An Agent Skill is often just a markdown file on the machine that tells the agent to read a file, call an API, or run a shell command. The recurring risk factors:

  • Inherited privilege. The plugin acts as the user, so it sees source code, credentials and internal systems the user can see.
  • No review gate. Components are installed by the employee, frequently from public registries, without a security approval step.
  • Instruction-level attack surface. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.
  • Interdependence. Agents, models, Skills, hooks, rules files and connectors interact, so assessing one component in isolation misses what the combination can do.

How is a risky plugin different from shadow AI?

Shadow AI means AI tools, agents, MCP servers and Skills running inside the organization that security never approved and cannot see — typically installed by employees on their own endpoints, often under personal accounts, such as an engineer connecting through a private Gmail login on a corporate machine. A rogue agent is the opposite situation: an approved agent taking actions nobody asked for, like reaching for credentials, escalating privilege or sending data to an unapproved destination, usually after manipulation or departure from its assigned task. A third-party plugin can produce either outcome — it may be unsanctioned software, or it may be sanctioned software that behaves beyond its brief. Backslash Security detects and prevents both patterns, including the use of private account access on enterprise machines.

Why doesn't the existing endpoint security stack catch it?

Endpoint detection and response (EDR) was built to catch malicious processes, files and known-bad behavior on a machine. A third-party AI plugin usually involves none of those: a signed agent binary reads a markdown file, calls a legitimate API over an approved connection, and writes to a repository the user already has rights to. Every individual step looks normal at the process boundary. What EDR does not observe is the cause: which prompt, agent, Skill or MCP server produced the action. Without the cause, the action cannot be stopped in time. That instruction-and-action layer is the subject of agentic AI endpoint security, and it is the layer Backslash Security monitors on employee machines.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to get started?

See how Backslash Security can help.

Book a Demo