-
September 10, 2026
-
September 10, 2026

AI can accelerate reconnaissance, generate convincing social-engineering content, and assist with malicious code development. But what happens when AI is no longer just the weapon, and the AI workspace itself becomes the attack surface?
Employees from a medical technology company received invitations to join what looked like their employer's Claude environment. The invitations came through Anthropic's real infrastructure, pointed to the real Claude service, and used a real authentication flow. Everything about the delivery was legitimate.
But the organization issuing the invitations was not legitimate. An attacker had created a workspace with a name closely resembling the victim company and used it to pull corporate data into a tenant the company didn't control. The result - sensitive company data was voluntarily exfiltrated from employees’ endpoints to an attacker.
Here’s how the attack unfolded:
The attacker creates an organization in the target AI service (in this case, Claude) and uses a lookalike name or a homoglyph attack: a display name that resembles the victim. The variation may use a duplicated character, punctuation change, abbreviation, subsidiary name, or other difference that is easy to overlook in an account selector or invitation screen.
Policies that check only domain or display name can't tell the victim's tenant from the attacker's, since the identifier that matters is the immutable tenant ID assigned by the provider.
The attacker sources employee emails from company pages, LinkedIn, conference materials, breach data, or contact databases, then uses the provider's normal invite flow. This breaks several trust signals the organization relied on:
None of that confirms that the inviting organization is the employee's company, only that the provider sent the invite.
The recipient authenticates normally and accepts the invitation, and the employee's account joins the attacker's organization.
The employee’s identity is genuine, but the tenant context is wrong. However, controls that record only the application name, destination domain, process, or successful authentication event might classify the session as benign.
If the employee begins using the workspace for normal business tasks, they might submit sensitive information through prompts, uploaded files, connected applications, or coding workflows.
Potentially exposed information can include credentials, API keys, and access tokens, proprietary business documents, customer or patient information, internal product plans and intellectual property, source code and engineering documentation, and sensitive details about internal systems and processes.
This information could also enable lateral movement within the environment, as attackers use valid credentials to access the endpoint and potentially reach repositories, cloud resources, or internal applications. From there, the incident could lead to data exposure, unauthorized access, privilege escalation, or broader compromise across connected systems.
The attack depends on four conditions:
The attacker does not need to impersonate the AI provider. Instead, the attacker uses the provider’s normal account-provisioning and invitation mechanisms. This is similar to abuse of legitimate cloud tenants, OAuth applications, collaboration platforms, and file-sharing services: the service is authentic, but the administrative domain inside it is hostile.
The primary security failure is failure to bind an approved corporate identity and managed endpoint to an approved AI tenant.
Tenant impersonation represents the inbound threat, but the risk surface on the endpoint extends further. As agentic AI adoption accelerates, risk enters through unmanaged local components even without an active external adversary:
Endpoint management tools inventory installed apps. EDR sees processes and file activity. Secure web gateways enforce domain policy. None of them understand the agentic fabric itself: which agents, models, skills, hooks, rules, plugins, MCP servers, and credentials are actually operating on a given endpoint, their capabilities and their behavior.
Backslash Security gives organizations visibility and control over the AI agentic fabric operating across employee endpoints. It continuously discovers every aspect of that fabric, maps every agent's full context layer, analyzes every component and combination, and identifies risk.
Control sits in the same place as the visibility: Backslash enforces policy inside the coding agent itself, on the employee's own machine, at the moment an agent signs in, loads an instruction file, reaches for a tool or sends a prompt to a model, so whatever an outside workspace pushes into a session is judged against the organization's policy and not against its owner's. And it evaluates intent before and during runtime, measuring what the agent was asked to do against what it actually reaches for.
In this attack, Backslash would recognize that Claude itself is legitimate while detecting that the employee is attempting to join a workspace with a different tenant ID from the organization’s approved Claude environment. The platform correlates that mismatch with the lookalike organization name and treats the connection as a potential impersonation attack, rather than leaving it classified as unknown Shadow AI or trusted Claude activity.
Backslash can then block access to the unauthorized workspace at the invitation stage, before it becomes part of the employee’s working environment and before enterprise tools, conversations, files, or credentials are connected. This binds corporate users and managed endpoints to approved AI tenants, closing the trust gap that legitimate domains, authentication flows, and provider-generated invitations cannot address on their own.
Backslash Security is the Agentic AI Endpoint Security platform. We enable enterprises to discover, govern, and protect the agentic AI fabric - every AI agent, MCP server, and Skill running on employee endpoints - securing agentic AI at enterprise scale and business velocity.