FAQ

Why AGENTS.md Files Belong in Your AI Vetting Process

At a glance

  • A rules file such as AGENTS.md carries standing instructions an AI agent reads on every run, executing with the employee's own identity and access.
  • Backslash Security research found hidden instructions in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials and npm tokens.
  • Vetting a rules file means knowing it exists, who can edit it, what it instructs, and whether changes are recorded.
  • Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity.

Backslash Security

Published:

AGENTS.md files belong in your AI vetting process because an AI coding agent reads them on every run and acts on what they say. A rules file — AGENTS.md, CLAUDE.md and their equivalents — is a file in a repository or on an employee's machine that carries standing instructions shaping the agent's subsequent behavior, and those instructions execute with that employee's own identity, credentials and access. Anyone who can write to the repository can change what the agent does next, and where rules files sit outside the approval workflows that govern installed software, that change is invisible to security. Backslash Security research found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration — a concrete case of indirect prompt injection, in which hidden instructions planted in content an agent reads are followed as though the user had typed them.

That places rules files inside the scope of agentic AI endpoint security rather than general tool approval. Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader, a badge awarded to vendors demonstrating the most in-depth controls for AI on the endpoint, including permission mapping, approved commands, and runtime controls for MCPs and Skills. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.

What is an AGENTS.md file, and what does it actually control?

An AGENTS.md file is a plain-text instruction file — a rules file, in other words — that lives in a repository or on an endpoint and carries standing directions an AI coding agent reads on every run. The agent treats its contents much like instructions a user typed at the prompt. Variants exist under other filenames, including CLAUDE.md, and as of 2026 both the format and which assistants honor it are still shifting, so any inventory of readers is accurate only on the date it was taken.

What attributes of an AGENTS.md file matter for vetting?

Attribute Typical values Why it matters
Location Repository root, nested per-directory, or user home directory Determines whether one team or every project on the machine inherits it
Format Markdown prose, usually unsigned and unreviewed Anyone with write access can change agent behavior without a code review
Permissions guidance Approved and forbidden commands, auto-approve settings Can widen what the agent may execute under the employee's own identity
Tool references Named MCP servers, scripts, build and test commands Declares which external connections the agent may act through
Conventions Style rules, directory layout, do-not-touch paths Shapes output quality and the blast radius of a bad run

Which adjacent terms should be defined before you vet one?

  • MCP server (Model Context Protocol): a connection an agent acts through to reach external tools and data.
  • MCP tool: a single callable function exposed by such a server.
  • Skill: packaged instructions and scripts that extend an agent, executing with the user's own permissions.
  • Hook: a trigger that fires before or after an agent action, running something alongside it.
  • Plugin / connector: components that attach further capability or data sources to the agent.
  • Context file: any file whose contents enter the agent's working instructions, of which a rules file is one kind.

Assistants in this class — among them Claude Code, Cursor and Codex — pick up instruction files from the paths they are configured to read, which is why the same file can steer a build command on one machine and a credential-touching command on another. Together these components form the agentic AI fabric on an endpoint, and an instruction file can reference or influence each of them.

Why does a plain Markdown file deserve a place in AI vetting at all?

A plain Markdown file earns a place in AI vetting because the agent does not read it the way a person reads a document. A rules file — AGENTS.md, CLAUDE.md or an equivalent carrying standing instructions an agent loads on every run — is consumed before the agent acts, and its contents shape which tools get invoked, which shell commands run, and which external servers get contacted. Those actions execute with the employee's own identity, credentials and access. This means the file functions as policy that runs, even though nothing about its format says so.

That difference changes how it should be classified. A passive document sits in a repository and waits for a human to open it; it has no reach beyond the reader. An instruction surface reaches the tool layer. It can point an agent at an MCP server — the Model Context Protocol connection an agent acts through to reach external tools and data — pre-approve a command pattern, or direct the agent to read a directory it would otherwise leave alone. Nothing in the file announces this; the text looks identical either way, and the privilege comes from the agent's trust in it, not from the file's permissions.

What should you do about it, and what can go wrong?

Do this But watch out for Mitigation in practice
Inventory rules files across endpoints, not just in central repositories Copies live on individual machines and in forks, so a repository-only view is incomplete Discover rules files at the endpoint, where the agent actually reads them
Review rules files for instructions that grant tool access or pre-approve commands Review is point-in-time; a file edited after approval carries new instructions on the next run Apply continuous assessment and allowlisting rather than one-off sign-off
Treat untrusted repositories as a source of agent instructions Content the agent fetches on its own can carry hidden directives it follows as though a user had typed them Enforce inline controls on the resulting action — credential access, outbound destinations — before execution

Which risks hide inside an unreviewed AGENTS.md file?

The risks that hide inside an unreviewed AGENTS.md file sit at the action layer, not the code layer. A rules file — a file in a repository or on a machine that carries standing instructions an agent reads on every run — can quietly change what an agent is allowed to do with the user's own identity and access. Common patterns worth looking for include:

  • Autonomy broadening. Standing instructions that tell the agent to skip confirmations, run commands without asking, or continue until a task is "done."
  • Pointers to unvetted tooling. Lines that instruct the agent to connect to an MCP server or tool nobody assessed, where MCP, the Model Context Protocol, is how agents reach external tools and data.
  • Data-handling directives. Instructions to send files, context, or output to a third-party endpoint that was never approved for company data.
  • Hostile instruction text arriving in content. Hidden directives merged through a pull request or inherited from a template repository, which the agent may follow as though a user typed them.
  • Drift between reviewed and running files. The version a team signed off on may differ from the one sitting on the machine.
  • Duplicate or nested files. Multiple rules files at different directory levels can override one another in ways the author did not intend.

What makes these patterns consequential is the execution context: the agent runs with the employee's credentials, shell access and network reach, so an instruction it reads is effectively an instruction the user issued.

Do this But watch out for
Require review of rules files in pull requests Review catches the committed file, not the local copy — pair it with endpoint-side discovery of what is actually present
Allowlist which MCP servers a rules file may reference Allowlists age; treat them as policy to re-enforce continuously, not a one-time list
Flag instructions that remove human confirmation Overly broad flags create noise — scope them to actions like credential access and outbound data

Is this the same as scanning source code for bugs? No. The question here is what the agent is instructed to do and what it can reach, which is a governance and runtime control problem.

How do you build AGENTS.md review into an existing AI vetting workflow?

You can build AGENTS.md review into an existing AI vetting workflow by treating the file as a reviewable artifact with an owner, not as developer scratch. AGENTS.md is a rules file — a file in a repository or on a machine that carries standing instructions an agent reads on every run — and it sits alongside CLAUDE.md, hooks, Skills and MCP server configs. The stages below are rollout work for teams already evaluating controls, and they apply to every employee endpoint; developer workstations are the most exposed case, not the boundary.

What does each stage involve, and who owns it?

  1. Discover. Sweep all endpoints for rules files, hooks, Skills, plugins and MCP server entries, distributing the collection through the MDM you already run — Intune or Jamf. Owner: endpoint/MDM administration. Cadence: continuous.
  2. Inventory what each file references. Record the tool calls, MCP servers, scripts, file paths, credential stores and external destinations an instruction file points at. Owner: security architecture. Cadence: on every change.
  3. Assess intent and blast radius. Read what the file actually authorizes, then map it to the identity it inherits — the employee's own account, tokens and repository access. Owner: AI governance lead. Cadence: triage on discovery, review on a regular rhythm.
  4. Set an approval and change-control path. Allowlist approved components, denylist risky ones, and write per-team policies so a research group and a finance team are not held to one profile. Pair this with your existing change tickets in Jira or ServiceNow. Owner: security with the component's team. Cadence: per request.
  5. Enforce inline. Backslash Security blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations — so a poisoned instruction is stopped rather than reported afterward. Owner: security operations. Cadence: always on.
  6. Reconstruct after the fact. Retain the audit and tracing record of agent activity, and route it to your SIEM so an investigator can establish which instructions were in place when an action fired. Owner: incident response. Cadence: on incident.

Which review criteria separate a low-risk rules file from a risky one?

Review criteria that separate a low-risk rules file from a risky one are mostly about granted authority, not prose quality. A rules file — AGENTS.md, CLAUDE.md or a sibling context file — carries standing instructions an agent reads on every run, so it behaves less like documentation and more like configuration. Define the criteria before you compare files, because each one is decisive in a different situation: autonomy scope matters most where an agent can execute shell commands, provenance matters most where the file arrives with a cloned repository, and precedence matters most where several context files are in play at once.

Review criterion What a reviewer should look for What should trigger escalation
Scope of granted autonomy Named, bounded tasks; explicit stops before destructive or irreversible actions Blanket permission to run commands, install packages, or act "without asking"
External references Which MCP servers, tools and connectors the file points the agent toward References to unreviewed endpoints, or tool calls introduced without an owner
Data movement instructions Clear limits on what leaves the machine and where it goes Instructions to read credential stores, environment files, or post content outward
Provenance and authorship A known internal author and a reviewed commit history Files inherited with third-party repositories or edited by unidentified contributors
Change frequency A stable file with infrequent, reviewed edits Silent edits between runs, or changes that arrive with dependency updates
Override and precedence Which file wins when machine-level, repository-level and user instructions conflict Local overrides that quietly supersede an approved baseline

The mechanism behind the provenance and data-movement rows is worth stating plainly. An agent cannot distinguish a line its operator wrote from a line someone else placed in a file it happens to read, and whatever it then does runs with that employee's own identity, credentials and access. Review therefore weighs authorship and reachable destinations as heavily as wording.

Escalation should route to whoever owns the affected system. Backslash Security converts an approved baseline into enforced allowlist and denylist policy, so vetting decisions persist as continuous guardrails instead of living in a reviewer's memory.

Frequently Asked Questions

What is an AGENTS.md file, and why does it belong in AI vetting?

An AGENTS.md file is a rules file — a plain markdown file in a repository or on a machine that carries standing instructions an AI agent reads on every run — and it belongs in vetting because it silently shapes what the agent does with the user's own permissions. CLAUDE.md serves the same purpose for a different agent. Nothing about it looks dangerous: it is text, it is committed like any other file, and it is rarely reviewed with the care given to executable changes. Whoever can edit that file can change agent behavior on every machine that reads it.

How can instructions inside a markdown file actually cause an incident?

Through prompt injection: hidden instructions placed in content an agent reads, which the agent then follows as though an operator had typed them. Security research published by Backslash Security found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. The agent was not compromised in the traditional sense — it did what its instruction set told it to do. That is why the instruction layer, not only the binary, needs review before an agent is approved for use.

Does our existing endpoint stack already cover this?

Generally not at the instruction layer. Endpoint detection and response was built to catch malicious processes, files and known-bad behavior on a machine; it sees a sanctioned agent binary executing normally, not the standing instructions that told it what to do. Backslash Security blocks risky agent actions inline before execution — including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations — which is the control point a rules file's instructions ultimately reach. Latio's evaluation found that "where Backslash stood out in our evaluation is in control depth," beyond scanning for what is malicious.

What else should a vetting process cover besides rules files?

Everything the agent can reach, because these components interact and none is safe to assess alone. Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint. A few definitions worth fixing before you write policy:

Component What it is Why vetting misses it
Rules file (AGENTS.md, CLAUDE.md) Standing instructions read on every agent run Reviewed as documentation, not as control logic
Agent Skills Packaged instructions and scripts that extend an agent, executing with the user's permissions Often just a markdown file nobody has audited
MCP server A Model Context Protocol connection an agent acts through to reach tools and data Installed per-user, invisible to device management
Hook A trigger that fires before or after an agent action Lives in local configuration, not in any inventory

Which coding agents does this apply to in practice?

The widely deployed ones your engineers are already running. Backslash Security covers AI coding agents including Claude Code, Cursor, Codex, Devin, GitHub Copilot and Antigravity. Each reads local configuration and instruction files on startup, so the same vetting question applies across all of them rather than to one vendor's tool. This is also where shadow AI surfaces — unapproved agents, models, MCP servers or Skills installed by employees, frequently under personal accounts — because an unapproved agent arrives with its own instruction files attached.

How do we start without installing software on employee machines?

Begin with the free AI Endpoint Exposure Assessment from Backslash Security, which is agentless, read-only, and retains no data — a short script distributed through your existing device management that runs, reports and disappears. Sales conversations often refer to it as Scout. It gives you the inventory you currently lack before you commit to policy. For regulated environments, Backslash Security states that it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC, which is the record-keeping side of AI agent endpoint security that most teams discover they need only after an incident.


About this article

Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07

Still have questions?

Our team is happy to help.

Book a Demo