Blog

How to Set Per-Team AI Plugin Policies by Risk Profile

At a glance

  • A per-team AI plugin policy ties the strictness of agent extension rules to what each team's agents can actually reach.
  • Start with inventory: agents, MCP servers, Agent Skills, hooks, rules files and connectors rarely appear in standard device or process inventories.
  • Classify teams by blast radius — credentials, repositories, production systems, customer data — then allow, deny, or conditionally approve each component class.
  • Shadow AI, meaning unapproved agents and Skills installed under personal accounts, is what a per-team policy must be able to detect.
  • Backslash Security discovers these components on employee endpoints and enforces team-specific allowlists and denylists continuously.

Backslash Security

Published:

A per-team AI plugin policy by risk profile is a written, enforced rule set stating which AI agent extensions each team may install, run, and connect to, with the strictness of each rule tied to what that team's agents can reach. "Plugin" here covers the whole extension layer an agent loads at runtime: MCP servers, meaning connections an agent acts through to external tools and data under the Model Context Protocol; Agent Skills, which are packaged instructions and scripts that extend what an agent can do and execute with the user's own permissions; hooks, which are triggers that fire before or after an agent action; rules files such as AGENTS.md or CLAUDE.md, which carry standing instructions an agent reads on every run; plus connectors, models, and editor plugins in tools like Claude Code, Cursor, and GitHub Copilot.

Setting the policy follows a predictable sequence. Inventory what employees are actually running on their machines, since a Skill is often a markdown file in a folder and an MCP server is a line in a configuration file, so neither reliably surfaces in device management or endpoint detection and response (EDR) inventories built around installed applications and running processes. Then classify teams by blast radius: which credentials, repositories, production systems, and customer records their agents can touch. Map each component class to an allow, deny, or conditional decision per team. Finally, enforce at the moment of action, because an approved component can still drift or be manipulated mid-run through hidden instructions planted in content the agent reads.

Backslash Security operates at that enforcement point on employee endpoints. According to Backslash Security, the platform blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations — which is what turns a per-team policy from a document into a control. Organizations with no inventory to classify yet can begin with Backslash Security's free AI Endpoint Exposure Assessment, which is agentless — collecting information without leaving software permanently installed on the endpoint — read-only, and retains no data.

Why does one org-wide AI plugin allowlist break down across teams?

One org-wide allowlist for AI plugins starts to break down because the AI components it governs are not the same objects from team to team. This section narrows to a single sub-case: endpoint-installed extensions — plugins, connectors, Agent Skills (packaged instructions and scripts that extend what an agent can do), and MCP tools reached through Model Context Protocol servers, which are the connections an agent acts through. A global list names components. Policy decisions depend on attributes the list does not record.

Each approval actually turns on several attributes, each with its own range of values:

Attribute Values it can take Why it decides the policy
Component type Agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin, connector A Skill is often a markdown file on a laptop; a rules file such as AGENTS.md or CLAUDE.md carries standing instructions the agent reads on every run. Different objects, different review.
Host agent Claude Code, Cursor, GitHub Copilot, Codex, Devin, Antigravity The same connector behaves differently depending on the agent invoking it and the permissions that agent already holds.
Running identity Corporate SSO through Entra ID or Okta, or a personal account A personal login on an enterprise machine is shadow AI by definition, and no component-level allowlist detects it.
Data reach Source repositories, Git systems, cloud credentials, production endpoints, external destinations Marketing's document connector and an engineer's repository-wide MCP server carry unrelated blast radii.
Team risk profile Per-team policy, inherited default, exception with expiry Risk tolerance differs by function; one threshold forces either over-blocking or over-permitting.

Component reputation also shifts over time. Backslash Security's MCP Server Security Hub, a public and continuously updated risk database, held 81,021 MCP servers when read on September 22, 2026, each scored for risk, and a server's posture can change with an update long after approval was granted. An allowlist frozen at a point in time records a decision, while the attributes above keep moving underneath it.

What goes into a team's AI plugin risk profile?

What goes into a team's AI plugin risk profile depends on what you mean by "plugin," because the word carries two distinct meanings on an endpoint and each one produces a different policy.

The installed artifact. Here a plugin is a packaged extension an employee adds to an AI client: an MCP server (Model Context Protocol, the protocol agents use to connect to external tools and data sources), an Agent Skill — packaged instructions and scripts that extend what an agent can do — or an IDE extension in Cursor or Claude Code. A marketing team installing a community MCP server for a web-scraping tool is an artifact-level event.

The capability grant. Here a plugin is the set of permissions the artifact exercises when it runs under the employee's own identity. The same issue-tracker MCP server may hold read-only scope for a support team and a write-capable token reaching production infrastructure for a platform team.

This section uses both senses, with per-team profiles written against the capability grant, since an identical artifact carries different consequences depending on whose machine it runs on.

Which attributes define the profile?

  • Data sensitivity: what the team routinely touches — public marketing copy, customer records, source repositories, or credentials. Sets the ceiling on acceptable exposure.
  • Identity and permission scope: the accounts and tokens attached to the user, plus local credential stores the agent can reach.
  • Capability class: read, write, execute, or network egress — egress meaning data leaving the endpoint to an external destination. Execute and egress change the blast radius most sharply.
  • Publisher provenance: first-party, verified publisher, or anonymous community upload. Backslash Security's MCP Server Security Hub scores each listed MCP server for risk, which gives provenance a scoreable basis rather than a reputational guess.
  • Update cadence: auto-updating components can change behavior after approval; pinned versions do not.
  • Endpoint action scope: which shell commands, file paths, Git remotes and outbound destinations the agent may touch while running.

Which policy tier should each risk profile get?

Each risk profile should be mapped to a policy tier using the same small set of criteria, applied the same way to every team, so that two groups with different exposure end up in different tiers for reasons anyone can audit. Define the criteria before you compare the tiers:

  • Identity and blast radius. An Agent Skill — packaged instructions and scripts that extend what an agent can do — runs with the employee's own permissions. This criterion is decisive wherever the operator holds production or administrative access.
  • Data and credential reach. What the component can touch through an MCP server, the Model Context Protocol connection an agent acts through: source repositories, ticketing systems, cloud keys, customer records.
  • Provenance. Whether the agent, Skill, MCP server, hook or rules file comes from a vendor you have assessed, an internal build, or an unvetted public registry. This one decides most cases where the capability itself looks harmless.
  • Reversibility. Whether a mistaken action can be undone. Read-only retrieval and an outbound write to a third party sit at opposite ends.
Tier What qualifies Allowed capabilities Approval path Enforcement posture
Open Low reach, reversible actions, assessed provenance Read-only context, local file reads, search Self-service from the approved catalog Discovery and logging only
Reviewed Useful reach, assessed source, recoverable actions Repository reads, writes to non-production systems Owner request, security sign-off Logged, with alerting on deviation
Restricted Credential, production or customer-data reach Narrow, named tool calls only Named approver per team, time-bound Deny by default; approved actions permitted
Blocked Unassessed or known-risky components, personal-account connections None Exception only, with a documented owner Invocation refused

A platform or infrastructure team typically lands in Restricted on reach alone, while a marketing team using a research assistant with read-only sources fits Open on the same criteria. Backslash Security supports this split directly by allowlisting approved components, denylisting risky ones, and letting you write custom policies for different teams and risk profiles — so the same Skill can sit in Reviewed for one group and Blocked for another.

How do you enforce a per-team policy inline without stalling the team?

To enforce a per-team policy on the endpoint, the control point has to sit where the agent actually acts and has to know whose rules apply to the identity running it. A per-team policy here means an allowlist and denylist of agentic components — agents, MCP servers (the Model Context Protocol connections an agent acts through), Agent Skills, hooks, rules files and plugins — scoped to a group rather than to the whole company. This means enforcement cannot be a static inventory review: if the policy varies by team, the rule applied to an agent action has to be that team's rule, applied at the moment the action happens.

Do this But watch out for — and how to contain it
Discover what is installed before writing any rule A point-in-time audit ages immediately, since employees add components themselves. Contain it with continuous rediscovery, not an annual sweep.
Assess components, then govern them Blanket-denying anything unassessed stalls legitimate work. Start new teams on discovery and logging before blocking, then promote assessed components onto the allowlist.
Write a custom policy for each team and risk profile Team composition drifts, so a joiner can silently inherit a looser policy than their role needs. Review team assignments in the same cadence as the policies themselves.
Constrain actions rather than banning whole tools Tool-level bans push people onto personal accounts, which is how shadow AI starts. Permit the assistant, restrict what it may reach and where it may send data.
Act before execution on the highest-risk operations Over-broad interception creates visible friction. Keep real-time prevention narrow and route everything else to logging and review.
Give denied users a path back Unexplained blocks erode goodwill faster than any outage. Pair each denial with the trace and a named owner who can approve an exception.

Backslash Security governs this layer on the endpoint across teams and identities. Latio's evaluation found that "where Backslash stood out in our evaluation is in control depth" — control over endpoint agent settings and over what is available to an agent in the first place, rather than scanning alone.

How should you roll out, review, and prove per-team AI plugin policies?

Roll out per-team AI plugin policies in stages, and fix the review cadence before the first rule goes live. This is implementation-stage work: the controls already exist, and the task is switching them on in an order that gives each team a defensible policy without stalling the work it was hired to do.

A staged rollout sequence

  1. Inventory before you write a single rule. Discover what employees are already running — agents, models, MCP servers and tools, Skills, hooks, rules files, plugins and connectors. Agentless collection, meaning information gathered without leaving software permanently installed on the machine, gets you a baseline without a change-management fight.
  2. Segment by risk profile, not by convenience. Write a custom policy for each team and risk profile, so an engineering group running Claude Code or Cursor is governed separately from a finance group running a desktop assistant.
  3. Pilot with one team before enforcing. As a rollout practice, keep that team on discovery and logging first and watch real activity before switching on blocking. The pilot tells you which components are load-bearing and which are hobby installs.
  4. Allowlist, denylist, then tune. Approve what the pilot proved benign, deny what it did not, and record the exceptions with an owner and an expiry.
  5. Expand team by team. Distribute through the MDM you already operate, such as Intune or Jamf, and add one team's policy set at a time.

Review cadence

New components arrive constantly: a Skill that is often just a markdown file, a hook that fires before or after an agent action, a rules file such as AGENTS.md or CLAUDE.md carrying standing instructions an agent reads on every run. Review each team's policy set on a fixed schedule, and use continuous discovery as the trigger for an off-cycle review whenever a new component type appears on that team's endpoints.

Proving it afterward

Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome. That record answers the auditor's question from evidence: which prompt, which agent, which Skill or MCP server acted, and what it accessed.

Frequently Asked Questions

What is a per-team AI plugin policy?

A per-team AI plugin policy is a written, enforced rule set that says which agentic components each team may run on its endpoints, and which are denied. "Plugin" here is broader than a browser extension: it covers AI agents, models, MCP servers (Model Context Protocol connections through which an agent reaches external tools and data), MCP tools, Agent Skills (packaged instructions and scripts that extend what an agent can do, often just a markdown file on the machine), hooks (triggers that fire before or after an agent action), rules files such as AGENTS.md or CLAUDE.md that carry standing instructions an agent reads on every run, connectors, and plugins proper. Policy is normally expressed as an allowlist of approved components plus a denylist of risky ones.

How do you decide a team's risk profile?

Risk profile is a function of what the team's agents can reach, not of headcount. Useful criteria, defined before any allowlist is drafted:

Criterion What it describes When it becomes decisive
Identity and access Whose credentials the agent inherits when it acts Teams whose members hold production or admin access
Data in reach Repositories, ticket systems, customer records the agent can read Teams working with regulated or customer data
Write capability Whether agents can execute commands, commit, or deploy Engineering and platform teams
Component churn How often new Skills, MCP servers or rules files appear Fast-moving teams installing from public sources
Untrusted input exposure Whether agents read issues, web pages or third-party repositories Any team exposed to indirect prompt injection

Why can't MDM or the existing endpoint stack enforce this?

Device management platforms such as Intune and Jamf govern applications, configuration and device posture; a Skill or a rules file is a text file inside a user directory, invisible to that model of control. Endpoint detection and response is built to catch malicious processes, files and known-bad behavior, so an approved agent that reads a poisoned document and calls a tool with the employee's own permissions looks like ordinary activity. Enforcement has to sit at the agent action itself. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

How do you build an inventory before writing any policy?

You cannot allowlist what you have never enumerated, so discovery precedes policy. Backslash Security offers a free AI Endpoint Exposure Assessment — called Scout — that is agentless, meaning it collects information without leaving software permanently installed on the endpoint, read-only, and retains no data. It is distributed through existing device management tooling, runs, reports, and leaves. The output is the baseline for per-team rules: which agents, MCP servers, Skills, hooks and rules files already exist, and on whose machines.

What should policy say about Shadow AI and personal accounts?

Shadow AI is AI tooling, agents, MCP servers and Skills running inside the organization that security has not approved and cannot see — most concretely, an employee signing an enterprise coding agent into a personal account on a corporate machine. A per-team policy should treat account provenance as a first-class rule alongside component approval: enterprise identity through Entra ID or Okta permitted, private logins denied, with the denial enforced at the point of use rather than recorded in a document. Backslash Security detects Shadow AI, including private account access on enterprise machines, and blocks risky agent actions inline before execution.

How do per-team policies support audits and investigations?

Differentiated policy only holds up if you can show what each team was permitted to run and what actually ran. That means retaining the path of an agent run from prompt to agent to tool call to outcome, so an investigator can reconstruct which component acted, under which identity, and against which rule. According to Backslash Security, the platform traces that full path and automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. Latio's 2026 AI Security Market Report named Backslash an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, folder structures, approved commands, and runtime controls for MCPs and Skills — the control depth that per-team AI agent endpoint security policy depends on.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to get started?

See how Backslash Security can help.

Book a Demo