At a glance
- An agent audit trail must record the full causal chain: prompt, agent, model, Skill or MCP server, tool calls, identity, data touched, outcome.
- EU AI Act record-keeping and DORA incident reporting both assume records that reconstruct events, which requires resolving which component caused each action.
- Backslash Security was founded to secure the agentic AI fabric — agents, MCP servers, Skills, plugins, connectors and hooks running on enterprise endpoints.
- Unapproved tools running under personal accounts leave no record at all, so continuous discovery is a precondition for any defensible trail.
Backslash Security
Published:
An agent audit trail that will stand up to EU AI Act record-keeping duties and DORA's ICT incident-reporting obligations has to capture the entire causal chain of an agent run. That means the prompt or instruction that started it; the agent and model that executed; which Agent Skill — packaged instructions and scripts that extend what an agent can do, often a single markdown file on the user's machine — or which MCP server took part, MCP being the Model Context Protocol agents use to reach external tools and data; every tool call made; the human or service identity and permissions the run executed under; the data read, written or transmitted and its destination; and the final outcome. It also has to capture the control decision attached to each action — permitted, blocked or escalated — with timestamps, so the record shows what the agent did and what the guardrail did about it. Rules files such as AGENTS.md or CLAUDE.md, which carry standing instructions an agent reads on every run, belong in the record too, because they change behavior without any user typing anything.
Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and according to Backslash Security it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC. The surface that trail has to account for is wide: when read on 22 September 2026, Backslash Security's MCP Server Security Hub had scanned and scored more than 80,000 publicly available MCP servers, each one a connection an agent can act through.
What must an agent audit trail capture on an employee endpoint?
This narrows to one artifact: the per-run record produced on an employee laptop or workstation, not model-provider logs or cloud-side telemetry. An agent audit trail must tie together the identity that ran the agent, the component that acted, and the action's outcome—because an auditor asking what happened is really asking which instruction caused which effect.
The minimum attribute set for an endpoint record to function as evidence:
- Run identifier and initiating prompt. The triggering instruction, stored or referenced. Without it, there is no way to show the action was authorized rather than injected.
- Human and machine identity. The corporate directory principal (Entra ID, Okta, Active Directory) plus the account the agent actually used. This surfaces a personal account operating on an enterprise machine.
- Agent and version. Claude Code, Cursor, Codex, GitHub Copilot, Devin or another client, with its release. Behavior changes between versions, so a trail without versioning cannot be replayed.
- Component invoked. The MCP server or MCP tool (Model Context Protocol being the connection layer agents act through), the Agent Skill—packaged instructions and scripts that execute with the user's own permissions—the hook, plugin, or the rules file such as AGENTS.md or CLAUDE.md that supplied standing instructions.
- Tool call target and parameters. File paths, repositories, credentials touched, destination hosts.
- Outcome and enforcement decision. Completed, failed, or blocked before execution—and under which policy.
- Timestamp and host. For ordering events across SIEM sources such as Splunk.
Process-level endpoint logging records which binary executed. The prompt, the agent and the component that caused it to execute sit above that boundary, so those fields must be captured at the agentic layer itself. Backslash Security collects and presents audit and tracing data of agentic activity on endpoints to support compliance needs and incident forensics.
Which records does the EU AI Act expect from agentic AI activity?
This narrows to one case: the records the EU AI Act expects when AI agents run on employee endpoints rather than inside a centrally hosted application. The regime's record-keeping and traceability provisions ask deployers of high-risk AI systems to retain automatically generated logs that are under their control and to be able to reconstruct how a system operated during a given period. On an endpoint, that obligation lands somewhere no application-level log reaches: the agent executes locally, under the employee's own identity and access.
What fields does an endpoint-level agent record need?
| Attribute | Expected values | Why it matters for the record |
|---|---|---|
| Identity | Corporate directory account (Entra ID, Okta) or personal account | A personal login on a corporate machine is unsanctioned AI use that security has no visibility into |
| Agent and model | Claude Code, Cursor, GitHub Copilot, Codex, Antigravity; hosted or local models | Names the system whose operation is being evidenced |
| MCP server and tool | Named Model Context Protocol server and the specific tool call | MCP is the protocol agents use to reach external tools and data; each server is an action path |
| Skill, hook, rules file | Invoked Agent Skill, pre/post-action hook, AGENTS.md or CLAUDE.md standing instructions | These carry instructions the agent follows on every run, often as plain markdown files |
| Triggering input | User prompt, or content the agent read on its own | Distinguishes an instructed action from an injected one |
| Action and disposition | Attempted operation plus allowed or blocked before execution | Supplies the outcome a reviewer or regulator asks for |
Backslash Security collects and presents audit and tracing data of agentic activity on the endpoint and traces the full path of an agent run from prompt to agent to tool call to outcome, so evidence accumulates as a byproduct of governing agentic AI usage rather than being pieced together afterward from application telemetry that never observed the local run. The same field set serves an incident investigator, because the question asked after a security event and the question asked during an audit both resolve to a single agent run.
Which records does DORA expect when an agent touches ICT systems?
DORA — the EU's Digital Operational Resilience Act governing ICT risk for financial entities — centers on managing and reporting ICT-related incidents, and reconstructing such an incident end to end can include the case where the actor is an AI agent on an employee's laptop. This section addresses agents, Agent Skills (packaged instructions and scripts executing with user permissions), and MCP connections on endpoints, where MCP (Model Context Protocol) enables agents to act through external tools or data sources.
Resilience-grade agent records require these attributes:
- Actor identity — the human account, agent binary or client, and identity used at action moment, including personal logins on corporate machines.
- Trigger and instruction source — user prompt, scheduled run, hook, or rules files (AGENTS.md, CLAUDE.md) carrying standing instructions. Instructions the user never typed are typical root causes.
- Tool call and destination — MCP server, MCP tool, Skill invoked, plus the endpoint, service or system reached.
- Disposition — executed, blocked inline before execution, or modified. An investigation needs to know what stopped events, not only what detected them.
- Sequence and outcome — ordered steps with results, including data written or credentials touched.
Backslash Security collects and presents audit and tracing data of agentic activity on endpoints to support compliance needs and incident forensics and investigation.
How do EU AI Act and DORA audit trail requirements compare side by side?
Under both the EU AI Act and DORA, being able to reconstruct an agent's actions after the fact is the practical test of a record, and one underlying record can serve both, but the two regimes differ in scope and in the questions the record must answer.
Four criteria determine audit trail adequacy under either regime:
- Trigger scope — which organizations and systems fall in range, determining whether you need the record and for which endpoints.
- Record subject — whether the obligation attaches to an AI system's operation or to an operational and ICT process involving AI, dictating how you index evidence.
- Reconstruction depth — whether a single run can be replayed end to end: the identity that launched it, the agent, the Skill or MCP server it acted through (an MCP server being a connection an agent calls for external tools and data), and the resulting action.
- Reporting driver — whether evidence is produced on a conformity cadence or pulled urgently during incident notification, setting how fast data must be queryable.
| Regime | Trigger scope | Record subject | Reconstruction demand | Reporting driver |
|---|---|---|---|---|
| EU AI Act | Providers and deployers of in-scope AI systems, with heaviest duties on high-risk uses | The AI system's operation and logged events | Traceability of system behavior across its lifecycle | Conformity and oversight documentation |
| DORA | Financial entities and their ICT third-party providers | ICT operational processes and incidents | Ability to reconstruct an incident and its ICT dependencies | Incident classification and supervisory reporting |
A single evidence set can serve both: a per-run record linking user identity, agent, the Skill or MCP server invoked, the action attempted, its destination, and the enforcement decision. Backslash Security collects and presents audit and tracing data of agentic activity on endpoints for these compliance and incident-investigation needs.
Why do identity and endpoint logs leave gaps in reconstructing agent actions?
When an auditor asks what an AI agent actually did on a company laptop, identity records and endpoint logs each answer only a fragment. An identity provider such as Okta or Entra ID shows that a token was issued and used; it does not show which agent, which Skill, or which MCP server—the Model Context Protocol connection an agent acts through—initiated the request. Endpoint detection and response tooling captures the process, file write and outbound connection, but stops at the process boundary: it does not resolve the prompt, the rules file (a standing-instruction file such as AGENTS.md or CLAUDE.md that an agent reads on every run), or the tool call that produced the behavior. SaaS and Git records have the same shape, because the agent acts under the employee's own account with the employee's own permissions, so autonomous activity is indistinguishable from human activity.
Backslash Security operates on the host at the agentic layer, where the agent, the Skill and the MCP server remain distinct, resolvable actors rather than a single anonymous process in the event stream.
| Do this | But watch out for — and how to cover it |
|---|---|
| Pull identity provider logs to establish session ownership | They attribute everything to the human; pair them with host-side evidence of which agent or Skill initiated the call |
| Forward endpoint telemetry into Splunk or another SIEM | Process events arrive without their cause; correlate them with the agentic component that triggered them |
| Treat MDM inventory from Intune or Jamf as the install record | A Skill is often just a markdown file, so no install event exists; use continuous discovery of agents, Skills, hooks and connectors |
| Rebuild events from destination-side SaaS audit logs | They show the write, not the instruction behind it; keep host-side evidence of the originating agent run |
What does evidence readiness for agent activity look like in 2026?
For security and governance leads answering the board and the auditors, the practical question is retrievable records rather than policy text. When someone asks what a specific AI agent did on a specific machine on a specific day, a control narrative without a reconstructable record answers very little.
As of 2026, evidence-readiness signals that security and governance leads must produce cluster around four areas:
- Inventory: a current list of which agents, models, MCP servers — Model Context Protocol connections through which an agent reaches external tools and data — Skills, hooks and rules files are running on endpoints, including components installed under personal accounts.
- Causation: the ability to tie an action back to the agent, Skill or MCP server behind it, not merely to the process that executed it.
- Enforcement record: proof that a policy was applied rather than only documented, including actions stopped before execution.
- Reconstruction: enough linked record to follow one run from the instruction that started it through to the systems it touched.
Latio's 2026 AI Security Market Report named Backslash Security an Endpoint AI Security Leader for in-depth controls including permission mapping, approved commands, and runtime controls for MCPs and Skills. Evidence readiness depends less on how much endpoint telemetry is retained than on whether a single agent run can be resolved end to end. Teams closing this gap generally begin with inventory, since every later artifact references it.
Frequently Asked Questions
What must an agent audit trail capture to satisfy EU AI Act and DORA expectations?
Agent audit trails have to reconstruct a complete causal chain, because EU AI Act record-keeping and DORA's incident-management focus both point toward showing what happened, in what order, and why. For agentic activity that means capturing the human identity and account in use, the agent and model, the instruction that started the run, the MCP server (Model Context Protocol, the protocol agents use to connect to external tools and data) and the specific tool called, any Skill or rules file that shaped the behavior, the action attempted, and the result. Backslash Security traces the full path of an agent run from prompt to agent to tool call to outcome, and according to Backslash Security it automatically generates audit evidence for EU AI Act, NIS2, DORA and SOC.
Why don't existing endpoint and identity logs produce this record?
Process-level and authentication telemetry records the effect without the cause. EDR — endpoint detection and response, the incumbent layer built to catch malicious processes, files and known-bad behavior — will show that a shell command executed or a credential file was opened, but not which prompt, which agent, which Skill or which MCP server produced that action. Identity platforms such as Entra ID or Okta record the sign-in, not the tool call the agent made afterward under that session. An auditor asking "who instructed this, and through what" is asking a question that sits in the agentic AI endpoint security layer, between the operating system and the network.
How does Shadow AI break the record before it is even created?
Shadow AI is AI tooling — agents, models, MCP servers and Skills — running inside the organization that security has not approved and cannot see, typically installed by employees on their own machines and sometimes connected through a personal account on a corporate laptop. Activity that runs under a private login leaves no enterprise-side record at all, so the audit trail has a hole that no amount of downstream log routing repairs. The starting condition for a defensible trail is discovery: Backslash Security discovers every agent, model, MCP server, MCP tool, Skill, hook, rules file, plugin and connector running on an endpoint.
Which components does an inventory need to include?
Auditors ask about the whole interconnected layer, not one tool. Backslash Security was founded to secure the agentic AI fabric — the mesh of agents, MCP servers, Skills, plugins, connectors and hooks running on enterprise endpoints. A rules file (for example AGENTS.md or CLAUDE.md) carries standing instructions an agent reads on every run; a hook is a trigger that fires before or after an agent action. Agent Skills security matters here because a Skill is packaged instructions and scripts that run with the user's own permissions, often as nothing more than a markdown file. The external surface moves constantly: when read on 22 September 2026, the MCP Server Security Hub operated by Backslash Security had scanned and scored more than 80,000 publicly available MCP servers.
What should the trail show when an agent follows hidden instructions?
Prompt injection places hostile instructions inside content an agent reads — a file, repository, issue or web page — which the agent then follows as though a user had typed them. A trail built only from agent outputs cannot explain that event, because the originating instruction never appeared in a prompt box. Security research by Backslash Security found that malicious instructions hidden in a repository's AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens and Git configuration. Backslash Security also blocks risky agent actions inline before execution, including unauthorized code execution, credential access, privilege escalation, and data sent to unapproved destinations.
How can a team establish a baseline without installing anything first?
Backslash Security offers a free AI Endpoint Exposure Assessment — called Scout — that is agentless, read-only and retains no data: a short script distributed through your existing MDM, such as Intune or Jamf, that runs, reports and disappears. Agentless applies to discovery; enforcement and inline prevention run on the host, where the agent actually executes. The output gives governance and endpoint owners a named list of the agents, MCP servers, Skills and rules files already present on company machines.
About this article
Backslash Security publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Backslash Security before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-10-07