Data Exfiltration

Stop Credential Exposure at Its Agentic Source on the Endpoint

Backslash stops the malicious instructions, skills, hooks, and agent behaviors that lead to credential exposure and data exfiltration from the endpoint.

Agentic AI: A New Kind of Data Exfiltration Risk

Whether using excessive permissions or ones obtained through escalation, AI agents can autonomously access credentials, sensitive data and enterprise systems, creating new exfiltration paths through trusted MCP servers, connectors, and APIs that legacy security tools can’t see.
Credential Exposure in the Wild
Backslash Research - Codex
The Backslash Security research team found that at a certain time, malicious instructions hidden in a repository’s AGENTS.md file could trick OpenAI Codex into silently accessing AWS credentials, npm tokens, and Git configuration.
Backslash Research - Claude Code
The Backslash Security research team  discovered that malicious repository configurations can redirect Claude Code API requests to an attacker-controlled server, silently exposing the user’s sign-in token. The stolen credential could allow attackers to run AI workloads, access profile information, upload files, and initiate sessions under the victim’s identity.

Backslash Protects the Agentic Control Plane

Instead of waiting to classify the credentials as sensitive data being transferred, Backslash can identify and block the malicious skill, poisoned hook, unsafe instruction, or risky agent behavior enabling the transfer.
Discover

Identify every AI agent, skill, MCP server, connector, plugin, hook, rules file, and agent configuration operating across employee endpoints. Create a complete inventory of the operational context influencing agent behavior.
Assess

Continuously evaluate agentic components for malicious instructions and behaviors that could lead to data exfiltration. Identify skills that instruct agents to retrieve credentials, hooks that execute unsafe commands, MCP tools with unnecessary access, rules that redirect data, and combinations of components that create an exfiltration path.
Govern

Approve trusted skills, hooks, MCP servers, connectors, plugins, and configurations. Restrict unverified or malicious components before they can influence agent behavior. Apply policies based on the agent, user, component, permissions, origin, tools, commands, and observed intent
Block in Real Time

Detect and Block in Real Time

Detect commands and actions that could enable exfiltration, such as collecting environment variables, retrieving credentials, packaging source code, connecting to an unauthorized service, or instructing a tool to transmit enterprise information. Block the malicious skill, hook, instruction, or agent action before the workflow can reach the data or complete the transfer.
Investigate

Full forensic audit trail

Capture a forensic-ready audit trail connecting prompts, loaded skills, MCP interactions, tool calls, processes, and resulting actions.

What Pre-AI Era Security Controls Miss

Pre-AI era DLP tools and network-based gateways can miss much of data exposure and exfiltration attempts due to lack of visibility and context into agentic activity, and limited control over exfiltration paths and evasion techniques.
No Dedicated Agentic Endpoint Security
With
Visibility into agentic exfiltration paths
Discovery of malicious skills, hooks, and instructions
Context across agents, data, tools, and destinations
Control over MCP servers, connectors, and plugins
Protection of sensitive data collected or transferred
Investigation of agent-driven data movement
Interception of data before encoded network transfer and evasion attempts
Limited
Inconsistent
Fragmented
Limited
After-the-fact
Partial
Partial
Complete
Continuous
Connected
Granular
 Real Time
End to end
Real time
Backslash gives us full visibility and governance over our evolving agentic AI ecosystem, helps us triage what actually matters, and never gets in the way of velocity.
Chris Niggel, Head of Security
Turn the lights on
Move from "we think our developers are using DeepSeek" to a live map of every AI tool, model, MCP, and integration in use.
Eliminate Shadow AI
Policy enforcement moves from a document that developers ignore to a centralized control layer. Shadow AI drops to zero for governed tooling.
AI audit-ready
For EU AI Act, NIS2, DORA, and SOC 2 obligations, Backslash generates the evidence of AI governance controls and events tracing automatically.
Be the Dept. of YES
Developers and workforce users adopt AI tools freely, enabled with guardrails — achieving significant efficiency gains without the exposure.

Don’t Let Agent Instructions Become Exfiltration Paths

Stop malicious skills, poisoned hooks, unsafe commands, and manipulated agent behavior before they expose your credentials, source code, customer information, or proprietary business data.
See Backslash in Action

Common questions about Data Exfiltration

What is agentic data exfiltration?

It is data leaving the organization through an AI agent's own legitimate activity rather than through a recognizable attack. An agent with excessive or escalated permissions can autonomously read credentials, source code and enterprise data, then move them out through trusted MCP servers, connectors and APIs. Every component involved is approved, every call looks ordinary, and no malware is present.

How is this different from traditional data loss prevention?

DLP waits for the data to move, then tries to classify it as sensitive. That is the wrong moment and the wrong signal for agentic activity, because the transfer happens through sanctioned channels under a valid user identity. Backslash works a step earlier, identifying and blocking the malicious skill, poisoned hook, unsafe instruction or risky agent behavior that enables the transfer in the first place.

Can an AI agent leak credentials without being attacked?

Yes. Excessive permissions are enough on their own - an agent given broad file and network access will read credential files and reach external services as part of completing an ordinary task. Attacks make it faster, but they are not required. This is why permission scope matters more than threat detection here.

How do malicious instructions reach an agent?

Through the files and components it already trusts. Our research found that instructions hidden in a repository's AGENTS.md could cause OpenAI Codex to silently access AWS credentials, npm tokens and Git configuration. In a separate finding, malicious repository configurations redirected Claude Code API requests to an attacker-controlled server, exposing the user's sign-in token - a credential that allows running AI workloads, uploading files and starting sessions under the victim's identity.

What makes a stolen agent token worse than a stolen password?

Reach. A sign-in token for an AI tool can allow an attacker to run workloads, access profile information, upload files and initiate sessions as the victim - with the agent's own permissions behind each action. Nothing about that activity appears anomalous, because it is the same identity doing the same kinds of things it always does.

Why don't network gateways catch this?

Because most of the relevant activity never crosses the network in a form a gateway can inspect, and what does cross it is a valid call to a legitimate service. A gateway sees an authorized request to an approved destination. It has no visibility into which skill instructed the agent, which MCP tool was invoked, or whether the retrieval was part of the task the user asked for.

Which exfiltration behaviors can Backslash detect?

Commands and actions that enable a transfer: collecting environment variables, retrieving credentials, packaging source code, connecting to an unauthorized service, or instructing a tool to transmit enterprise information. The malicious skill, hook, instruction or agent action is blocked before the workflow reaches the data or completes the transfer.

Can combinations of safe components create an exfiltration path?

Yes, and this is the case per-component scanning misses entirely. A skill that retrieves credentials, a hook that executes unsafe commands, and an MCP tool with unnecessary access may each pass review individually while together forming a working exfiltration path. Backslash assesses components in combination, not only one at a time.

How do I know which agent components are already running on our endpoints?

You need an inventory of the operational context, not just the applications. That means every AI agent, skill, MCP server, connector, plugin, hook, rules file and agent configuration across employee endpoints - including anything installed under personal accounts or outside security oversight. Discovery is the first of the four stages, because nothing downstream is possible without it.

What evidence is available after a suspected exfiltration?

The chain connecting intent to outcome - which component introduced the behavior, what the agent accessed, which tools it called, and where data was sent. That sequence is what separates a malicious component from a legitimate but overscoped action, and it is the first question asked in any investigation.

‍

Github Copilot Logo Claude Logo Devin Desktop Logo Antigravity Logo Openclaw Logo Cursor Logo MCP Logo Gemini CLI Logo Codex Logo