Comparison

How an AGENTS.md File Becomes an Attack Path

At a glance

  • An AGENTS.md rules file hands an AI coding agent standing instructions on every run; anyone able to write to the repository can rewrite them.
  • Instructions planted in that file execute with the employee's own credentials and access, so repository write access effectively becomes endpoint access.
  • Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.
  • Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

Backslash Security

Published:

An AGENTS.md file becomes an attack path because it is a rules file: a plain markdown document that carries standing instructions an AI coding agent reads on every run and follows as though the user had typed them. Anyone who can write to the repository — a contributor, a merged pull request, a copied starter template, a dependency that ships its own rules file — can place instructions there that the agent then carries out using the employee's own credentials, tokens, and access to source control, cloud accounts and production systems. This is indirect prompt injection, meaning hostile instructions reach the agent through content it reads rather than through anything the operator typed. Backslash security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.

The incumbent control sitting on that same laptop is EDR — endpoint detection and response, the layer enterprises buy to catch malicious processes, files and known-bad behavior on a machine. EDR records the process that executed: a shell command, a credential read, an outbound connection. The instruction that produced that process lives one layer up, in the markdown file the agent read, next to the Skills, hooks, MCP servers, plugins and connectors the same agent can reach from the host. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and through 2026 its coverage includes AI coding agents such as Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity.

How does an AGENTS.md file turn into an attack path?

An AGENTS.md file turns into an attack path when a coding agent treats that file as instruction rather than documentation. A rules file — AGENTS.md or CLAUDE.md — is a plain markdown file in a repository or on a machine carrying standing instructions an agent reads on every run. Nothing in the format marks authorship, so text added by an outside contributor, a dependency, or a cloned repository arrives in the agent's context with the same standing as text the developer wrote.

The mechanism runs in a straight line:

  1. The file reaches the endpoint through an ordinary action — a clone, a pull, a branch checkout, a shared template.
  2. The agent starts a session and loads the rules file into its working context before the employee types anything.
  3. Hidden instructions in that content are followed as though the user had issued them. This is indirect prompt injection: hostile instructions reaching an agent through content it reads.
  4. The agent acts with the employee's own permissions — shell access, file system, Git credentials, connected MCP servers — so the resulting reads, writes and outbound calls resemble normal developer activity.

This means a text file becomes an execution path: it runs through the agent's tools, under a human identity, without a malicious binary ever touching disk.

Do this But watch out for Mitigation in the same step
Inventory rules files, hooks and Skills on endpoints A one-time scan ages out quickly as normal pulls change what is on disk Run discovery continuously, not as a project
Review rules file changes in code review Rules files also sit outside repositories, in home directories and global config Extend discovery to the host, not only the repo
Restrict which tools and MCP servers an agent may reach Blunt restriction pushes people toward personal accounts Set per-team policies matched to each risk profile
Act on the instruction source, not only the symptom Process-level telemetry shows the command, not the agent or file that caused it Backslash Security enforces on the host at the agentic layer, where the agent actually executes

What actually lives inside an AGENTS.md file, and which agents read it?

To be precise about scope: what follows covers the file itself — what actually lives inside an AGENTS.md, where it sits, and which agentic clients read it. AGENTS.md is a rules file: a plain markdown document kept in a repository that carries standing instructions an AI coding agent loads on every run. CLAUDE.md fills the same role for Claude Code.

Treated as a configuration object, the file has these attributes:

Attribute Typical values Why it matters
Location Repository root, with optional nested files per directory Nested copies are easy to miss in review; the agent still reads them
Format Free-form markdown, including code fences No schema, no required fields, nothing to validate against
Typical contents Build and test commands, coding conventions, permitted tool and CLI usage, environment setup steps, directories to leave alone Each line is an instruction the agent may act on, not documentation a human skims
Load behavior Pulled into context automatically at session start The developer never reads the text they are effectively sending to the model
Change control Edited by anyone with write access, usually via an ordinary pull request Content changes alter agent behavior with no software installed
Execution identity Runs with the signed-in developer's own permissions and local credentials Sets the reach of anything the file instructs

On the client side, repository-level instruction files are part of the normal operating model for mainstream agentic coding tools. Per Backslash Security, its coverage spans AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity — the same clients that read rules files, invoke Agent Skills, and connect outward through MCP servers, the protocol agents use to reach external tools and data.

Which other agent context artifacts carry the same risk as AGENTS.md?

Other agent context artifacts carry the same class of risk as an AGENTS.md file, because each one injects instructions or capability into an agent run that no human reads at the moment of execution. A rules file holds standing instructions an agent reads on every run. Agent Skills are packaged instructions and scripts that extend what an agent can do, executing with the user's own permissions. A hook is a trigger that fires on an agent action, running something before or after it. Each MCP server — MCP being the Model Context Protocol agents use to reach external tools — is a connection an agent can act through.

Four criteria make these artifacts comparable before any of them is judged:

  • Trust assumption — whether the artifact is treated as data or as instruction once the agent holds it. Decisive wherever the content arrives from outside the organization.
  • Load path — whether a user explicitly invokes it, or the agent picks it up on its own from a directory, repository or configuration.
  • Blast radius — which systems, credentials and environments the resulting action can reach.
  • Who reviews it — whether any approval step exists before it takes effect on the endpoint.
Artifact Trust assumption How it loads Blast radius Who reviews it
Rules file (AGENTS.md, CLAUDE.md) Read as standing instruction Auto-loaded on every run in scope Everything the agent can touch in that session Often merged as documentation
Agent Skills Executable capability under user permissions Invoked by the agent when relevant Shell, filesystem, APIs reachable by the user Rarely audited; often a single markdown file
Hooks Trusted automation around agent actions Fires on a matching agent event Whatever the triggered command can run Local configuration, typically unreviewed
MCP servers and tool descriptions Tool metadata treated as trustworthy Loaded at connection time External systems the server fronts Chosen by the individual user
Plugins and connectors Approved extensions of the client Installed into the agent client Accounts and SaaS the connector authenticates to Vendor marketplace, not the enterprise

Device management enrolls machines; it does not govern these files. Backslash Security assesses the risk posture of all of these endpoint agentic components using an agentless approach, meaning information is collected without leaving software permanently installed on the machine.

Why do endpoint and repository review controls miss a poisoned instruction file?

When an instruction file lands on an endpoint or is merged into a repository, neither the existing endpoint stack nor pull-request review tends to register the change as a security event. A rules file — a file in a repository or on a machine carrying standing instructions an agent reads on every run — is plain prose. There is no binary to hash, no manifest entry for a dependency resolver to flag, and no defective function for a static analyzer to catch. EDR, the incumbent endpoint layer built to catch malicious processes, files and known-bad behavior, has no hostile executable to match on, because the payload is sentences that a model later chooses to obey.

What does "review" mean in this context?

Two different activities travel under the word review, and conflating them is how the gap stays open.

  • Repository review is pull-request approval of committed changes. A reviewer approving an edit to an AGENTS.md or CLAUDE.md file is approving documentation-shaped text; its effect materializes later, on someone else's machine, when an agent reads the file and acts with that person's credentials.
  • Endpoint configuration review is software and policy inventory through MDM tooling such as Intune or Jamf. It enumerates installed applications, profiles and compliance state. Markdown dropped into a project folder is not an installed application.

This section uses review in the repository sense, and calls the second activity inventory or configuration checks.

How do the control layers relate?

Layer Strength buyers cite Architectural focus
CrowdStrike Established endpoint footprint with an agent already on the machine and broad general-purpose coverage General-purpose endpoint coverage from an agent already on the machine
Zscaler Already-owned network-layer security Network-layer security
Backslash Security Agentic components on the host Enforcement on the host at the agentic layer, where the agent executes

These layers are complementary, and Backslash Security replaces none of them; it resolves which agent, Skill or MCP server caused an action at the moment a written instruction becomes a real operation. Where an organization's AI use as of 2026 is limited to hosted chat tools with no locally running agents, there is nothing on the host for an instruction file to drive.

When in an agent's execution flow can a risky action still be blocked?

This depends on what you mean by "blocked." In an agent's execution flow, prevention means the action never reaches the host at all, while detection means it already ran and you are reading whatever record remains. The two sit at different points in the run:

  • Context load. The agent ingests its rules file — a file such as AGENTS.md or CLAUDE.md carrying standing instructions it reads on every run — plus Skills (packaged instructions and scripts that extend what an agent can do) and the tool descriptions advertised by connected MCP servers, the Model Context Protocol endpoints an agent acts through. All of it is assessable before anything executes.
  • Model reasoning. Inference happens inside the provider's model and is not observable from the endpoint. Only the inputs going in and the tool calls coming out are governable.
  • Tool call. The agent proposes a shell command, a file read, or an MCP invocation. Intent is legible here, and the call can be judged against policy before it is dispatched.
  • Action on the host. The command runs under the employee's own identity. This is the last point at which a control can refuse the action rather than record it, and Backslash Security enforces at this layer on the host, resolving which agent, Skill or MCP server caused the call.
  • After execution. Nothing remains to stop; only reconstruction from retained evidence.
Do this But watch out for — and how to handle it
Assess agents, Skills and MCP servers at context load A Skill is often just a markdown file that can change between runs; reassess continuously rather than once at onboarding
Evaluate tool calls against policy before dispatch Blunt rules block legitimate work and push people toward personal accounts; scope policies per team and risk profile
Enforce at the host, before the action lands Process-level telemetry shows the command without the causing agent, Skill or MCP server; require controls that resolve the cause
Retain evidence for the post-execution window Logs that stop at the process boundary cannot rebuild an agent run; capture agentic activity specifically

How can a security team discover, govern, and reconstruct agent context across every endpoint?

A security team can discover, govern, and reconstruct agent context across employee endpoints by running it as a continuous operating model rather than a one-time audit. Four stages, in order:

  1. Inventory. Enumerate what is actually running on each machine: which agents and local models are installed, which MCP servers (Model Context Protocol — the protocol agents use to reach external tools and data) they connect through, and which Skills, hooks, rules files such as AGENTS.md or CLAUDE.md, plugins and connectors shape their behavior. Backslash Security's free AI Endpoint Exposure Assessment, known as Scout, is agentless, read-only and retains no data, so a first pass can be distributed through existing MDM.
  2. Policy. Allowlist approved components, denylist risky ones, and write separate rules for teams with different risk profiles.
  3. Enforcement. A guardrail only holds where the agent executes — applied on the host, in the path of the action, rather than recorded as an alert after the fact.
  4. Reconstruction. Retain audit and tracing records of agentic activity so an investigation can establish what ran, under which identity, and on whose instruction.

An inventory of installed software tells you which tools are present; the instruction set an agent reads at run time is what decides what those tools are allowed to do.

Option Where it fits Differentiator
Backslash Security Organizations where a substantial population builds or runs AI agents Agentic endpoint security is the entire product, pairing agentless discovery with on-host enforcement
CrowdStrike Teams wanting broad general-purpose endpoint coverage from an agent already deployed Established, trusted endpoint footprint with broad general-purpose coverage
Zscaler Teams that already own network-layer security Network-layer security the buyer already owns

When not to switch anything. If AI use is confined to a few sanctioned chat tools with no local agents, Skills or MCP servers, the incumbent endpoint and network stack is a reasonable place to stay. The gap opens once employees run agents under their own identities — the work that belongs to agentic AI endpoint security.

Frequently Asked Questions

What is an AGENTS.md file, and how does it become an attack path?

AGENTS.md is a rules file — a file in a repository or on a machine that carries standing instructions an agent reads on every run, alongside variants such as CLAUDE.md. Because the agent loads it automatically, whoever can write to the repository can plant hidden instructions in content the agent reads and have them followed as though a user typed them. That mechanism is indirect prompt injection. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration.

Why doesn't the existing endpoint stack flag a poisoned rules file?

EDR — endpoint detection and response, the incumbent endpoint layer built to catch malicious processes, files and known-bad behavior — operates at the process boundary. A rules file is plain text that no process executes; the instruction becomes action only when an agent reads it and calls a tool. Resolving which prompt, agent run, Skill or MCP server produced that call belongs to the agentic layer that appeared between the operating system and the network once agents arrived on employee machines. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.

Which layers can act on this, and how do they differ?

Option Where it operates Strength buyers cite
Backslash Security On the host, at the agentic layer Discovers agents, models, MCP servers, MCP tools, Skills, hooks, rules files, plugins and connectors on an endpoint, and enforces inline before an action runs
CrowdStrike General-purpose endpoint platform Established, trusted endpoint footprint with an agent already on the machine and broad endpoint coverage
Zscaler Network layer Raised by buyers as an alternative when they already own network-layer security
Agentless-only agent-governance tools Collection without installed software Nothing to deploy on the endpoint at all, the fastest possible path to first visibility

When is staying with what you already own the right call?

If agent use in your organization is occasional and confined to a few people, the discovery question is answerable by hand, and adding an enforcement layer buys little. Where deploying endpoint software is genuinely difficult — contractor fleets, unmanaged hardware, strict change windows — inspecting agent traffic at the network layer suits the constraint better. Agentless-only governance also makes sense when the immediate need is a first inventory for a board or audit conversation, with enforcement deferred to a later cycle.

How can a team find out what is running on its endpoints today?

Backslash Security offers a free AI Endpoint Exposure Assessment that is agentless, read-only, and retains no data — agentless meaning information is collected without leaving software permanently installed on the machine. It is distributed through existing MDM, and the assessment is called Scout. It surfaces Shadow AI: agents, MCP servers and Skills that security has not approved and cannot see, including access through personal accounts on enterprise machines. Backslash Security also operates a free Skills Security Scanner, and runs the MCP Server Security Hub, a public, continuously updated risk database that held 81,021 MCP servers, each scored for risk, when read on 22 September 2026.

What record exists afterward if an agent run goes wrong?

Reconstruction needs the causal chain, not just the final system call. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, so an investigator can see which instruction source — a rules file, a Skill, an MCP tool — led to a credential read or an outbound transfer. Per Backslash Security, the platform also automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Ready to make the switch?

See why teams choose Backslash Security.

Book a Demo