Data Loss Prevention was built to answer one question: did something matching a known pattern just leave the perimeter? Credit card numbers, Social Security formats, labeled documents, keywords. It watches for the shape of sensitive data and blocks the match.
That model has a blind spot the size of your AI adoption curve. When an engineer pastes a block of source code into a personal ChatGPT account, there’s no pattern to match. When a sales rep drops a customer list into an AI browser extension, the content scanner sees plain text and lets a crown jewel walk out the door. When an AI agent reads a repository through an MCP server and summarizes it into an external model, no file crosses the boundary at all. The data left. DLP never woke up.
The scale is not theoretical. The Verizon 2025 DBIR found that 15% of employees routinely access generative AI on corporate devices, and that most of them do so through personal accounts that sit outside corporate authentication and logging. IBM’s 2025 Cost of a Data Breach report found that 63% of organizations still have no AI governance policy to control what employees feed into these tools, and Gartner predicts that by 2030, more than 40% of enterprises will hit a security or compliance incident tied to unauthorized shadow AI. The data is already moving, and most organizations can’t see it go.
This is the DLP problem restated for the AI era. Here’s why the old model can’t read the new exit, and what does.
What the Content Scanner Can’t Read
DLP works when data has a recognizable shape. A card number looks like a card number. A tagged HR document carries its label. The scanner matches the shape and acts.
Your most valuable data has no shape a scanner can match. A block of plain text might be a public blog post or the core of your pricing engine. A binary file might be a stock photo or a proprietary CAD design. A grid of numbers might be a spreadsheet or the weights of a trained model. Content inspection reads the bytes. It has no way to know that this particular block of text is the source code that took three years to write.
The categories DLP struggles to recognize are exactly the ones that matter most:
- Source code, which reads as ordinary text to a scanner
- Design and CAD files, which carry no keyword to match
- Model weights and training data, which look like undifferentiated numbers
- Board decks and strategy documents, whose sensitivity is contextual, not lexical
- Recorded meetings and transcripts, increasingly captured by AI notetakers
- Customer and HR records, which sometimes match a pattern and often don’t
When these move into an AI tool, DLP is reading for a pattern the data was never going to have.
Data type is only half of what a scanner misses. The other half is your business context, the rules that decide whether a specific movement is allowed in the first place. A finance team sharing an approved report with a customer is doing exactly what it should. The same file, pasted into a personal AI account by the same person, is a problem. The bytes are identical in both cases, so a content scanner treats them the same. It has no model of what your business actually permits, who is allowed to share what, with which customers, through which channels.
The New Exits DLP Was Never Watching
Even where DLP can match a pattern, AI opened data paths that route around the places DLP watches.
The prompt is the exit. A file upload to a corporate SaaS app crosses a boundary DLP monitors. A paste into a browser tab pointed at a personal AI account does not look like a file leaving at all. The most common way sensitive data reaches a public AI tool is an employee typing or pasting it into a chat window on a personal login, which means the text never passes a corporate control point on its way out.
The browser is the channel. Netskope reports that 72% of enterprise generative AI users access these tools through personal accounts, and that data flowing into AI apps grew more than thirtyfold in a year. Most of that traffic never touches an inline DLP proxy, because it originates in a browser session that authenticates personally.
MCP moves data with no file event. When an employee wires an AI assistant to an internal system through an MCP server, the agent can read from a database and pass the contents into an external model without a single file ever being downloaded. There is no attachment, no upload, no object for DLP to inspect.
The agent acts inside authorized access. An AI agent operating under a service account pulls data through an API it’s permitted to use. No login event, no MFA challenge, no anomaly a perimeter tool would flag. IBM’s Cost of a Data Breach report found that shadow AI added an average of $670,000 to breach costs, and Gartner projects that by 2027, 40% of AI-related data breaches will stem from cross-border generative AI misuse, meaning data pasted into a tool ends up processed in a jurisdiction your compliance program never mapped.
DLP was built for a world where data leaves as a file through a monitored door. AI data leaves as a prompt, a session, or an agent’s API call, through doors DLP was never standing at.
Identify Data by Where It Came From
If you can’t recognize sensitive data by its shape, you have to recognize it by its history. Provenance is the shift: instead of asking “does this content match a rule,” you ask “where did this data come from, what path did it take, and whose hands moved it.”
Provenance tracks three things a content scanner can’t:
- Origin. That block of text came from somewhere specific: a clone of a named repository, a row set from a customer table, an export from Salesforce. The scanner sees characters. Provenance sees where they came from.
- Path. The dataset moved from an internal store, through a local file, into a browser tab pointed at an AI service. Each hop is a fact, and the sequence is the story.
- Hands. Every step ties back to a person or an agent. Not “a file was uploaded,” but “this employee moved this data, which originated here, into this tool.”
A content scanner asks whether the bytes look sensitive. Provenance already knows the data is sensitive because it watched where it came from, and it can tell you whether this particular movement, by this particular person, looks anything like normal.
Same Paste, Three Stories
Consider a single event: an employee pastes a large block of source code into an AI tool. Without context, this is either an alert or, more likely, nothing at all, because to a content scanner it’s just text. With context, it’s one of several completely different stories.
Story one. The employee submitted their resignation on Monday. The code came from a repository outside their team’s scope, cloned yesterday for the first time. Their peer group has never used this tool. This is exfiltration staged as productivity. Escalate.
Story two. The employee is a senior engineer debugging a function they own, pasting it into an approved, corporate-authenticated assistant to find an error. The code is theirs, the tool is sanctioned, the pattern matches how their whole team works. This is Tuesday. No action needed.
Story three. No person pasted anything. An AI agent, connected through an MCP server the employee set up last week, read the repository and passed it to an external model on its own. No upload, no session, no human at the keyboard. This is an over-permissioned agent operating outside its scope. Constrain it and find out who wired it in.
One motion, three stories, three response playbooks. A DLP rule that fires on “large text block to unknown domain” can’t tell them apart, and a rule that stays quiet misses all three. The difference lives in provenance, identity, employment status, and peer-group history, correlated in one place.
How Anzenna Helps
Anzenna protects data flowing into and out of AI by watching lineage and identity instead of scanning content.
- Provenance over pattern matching. Anzenna reconstructs the journey of a sensitive dataset, from origin (a repository clone, a database row set, a Salesforce export) through every system it crosses, to the hands that moved it. It recognizes source code, design files, and model weights as what they are, not as text a scanner failed to flag.
- The AI exits, watched. Anzenna sees the paths DLP misses: browser paste and upload into AI services, personal versus corporate account use, OAuth data grants, and MCP server connections moving data with no file event. It has direct hooks into major AI platforms, including Claude, GitHub Copilot, Cursor, Gemini, Codex, and Windsurf, so it can see what those tools actually do on a managed device.
- Every movement tied to a person. A dataset leaving is noise on its own. A dataset leaving, moved by an employee in their notice period, whose peer group has never touched that tool, whose volume is well above their own history, is an investigation. Anzenna connects every hop to the human identity behind it, including employment status, role, and behavioral history.
- Peer-group baselines over blanket rules. Anzenna doesn’t block “code into an AI tool” for everyone, which would break the engineers who use it correctly all day. It flags the movement that deviates from how a person’s peers actually work. The baseline adapts to role, department, and individual history, so the alert fires on the story that matters rather than on the tool involved.
- Investigation Agents write the case. When an AI data risk surfaces, Anzenna’s Investigation Agents draft the case file automatically: what data moved, where it came from, which tool received it, who was behind it, how the behavior compares to baseline, and the recommended action. Your analyst reviews a narrative, not a raw log. Median case draft time is under 2 minutes.
- Agentless, read-only, metadata-only. Anzenna doesn’t inspect the content of what employees type into AI tools. It reads metadata: which tools, which data stores, how much, how often, by whom. Secrets and sensitive content are redacted before they enter the platform. No inline proxy, no endpoint agent, no change to your stack. SOC 2 Type II certified, 15-minute install.
Across its customer base, Anzenna has blocked 81,400 AI uploads and 79,500 exfiltrations, protecting more than 756,000 users. As one customer described it: “Anzenna caught one of our engineers uploading sensitive source code into their personal GitHub repo.”
Bringing It All Together
Traditional DLP asks whether the bytes look dangerous. That question stopped being enough the moment your most valuable data started leaving as a prompt, a browser session, or an agent’s API call, none of which carry a pattern to match. Protecting data in the AI era means knowing where it came from, watching where it goes, and reading the movement against the person making it.
The risk lives in the story behind the movement, and that story leaves whether or not a scanner recognizes the object walking out with it.
Ready to see how your data actually moves into AI? Request a demo and see it on your data in 30 minutes.
Statistics sourced from the Verizon 2025 Data Breach Investigations Report, the IBM 2025 Cost of a Data Breach Report, Gartner AI security predictions, and the Netskope Cloud and Threat Report 2025. Anzenna product metrics from anzenna.ai.
Frequently asked questions
What is AI DLP?
Data loss prevention aimed at AI data flows: catching sensitive content heading into copilots, chatbots and agents, including personal accounts, that classic pattern-based DLP misses.
Do AI DLP tools understand data context?
Context-aware AI DLP reasons over behavior, identity and where data is going, not just string patterns. Anzenna correlates signals across your stack so it understands intent, which is what pattern matching cannot do.
Why does traditional DLP fail on AI data flows?
Rule-based DLP matches known patterns. AI data flows, a paste into a chatbot or a prompt that pulls from a repo, often have no fixed pattern, so they slip past. Behavioral context catches them.
Is Anzenna's AI DLP agentless?
Yes. Anzenna connects through APIs across identity, SaaS, cloud and endpoint, with nothing to install.
Related reading: AI DLP, data loss prevention.