The lethal trifecta — why your agent is one email away from exfiltration
An LLM agent becomes dangerous the moment it holds three things at once: private data, exposure to untrusted content, and a way to talk to the outside world. Here is why that combination leaks, and what to actually remove.
A chatbot that can only talk is a low-stakes system. Give it your inbox, a browser, and an HTTP tool, and you have built something that can be instructed by a stranger and will act on those instructions with your credentials. Nothing was hacked. The model did exactly what it was designed to do.
The failure has a shape, and once you can see the shape you can find it in almost any agent architecture in about ten minutes.
The three ingredients
An agent is exposed when it simultaneously has:
- Access to private data — your email, your repos, your customer records, your internal wiki. Anything the attacker cannot read directly.
- Exposure to untrusted content — a web page it fetches, an email it summarizes, a PDF a user uploads, an issue comment, a calendar invite. Any text that arrived from someone who is not the operator.
- A way to communicate externally — an HTTP tool, an email send action, a webhook, a rendered image with a remote URL, even a Markdown link the user might click.
Hold all three and the attack writes itself: untrusted content carries instructions, the model follows them, private data leaves through the exit.
The model has no reliable way to tell your instructions from the data it is reading. They arrive in the same context window, in the same language, with the same authority. "Ignore the above" is not a bug in the prompt — it is a consequence of the architecture.
Why guardrails do not close it
The instinct is to filter. Scan the incoming content for injection attempts, add a system prompt that says never follow instructions found in documents, maybe run a classifier over the output.
These help at the margin and none of them are a boundary. The reason is counting: a filter has to catch every phrasing in every language and encoding, forever, while the attacker only needs one that works. Base64 the payload. Put it in alt text. Split it across two documents the agent will read in the same session. Write it as a plausible-looking system notice. The asymmetry never moves in the defender's favour.
Worse, a filter that mostly works is more dangerous than no filter, because it converts a known-open door into one everybody believes is shut.
Remove an ingredient instead
The trifecta is a conjunction, so you break it by deleting a term, not by inspecting one harder.
| Remove | How it looks in practice | What it costs |
|---|---|---|
| The exit | No outbound HTTP from the agent's tool set; allowlist the two domains it genuinely needs | The agent cannot "just fetch" things |
| The private data | Give the agent a scoped, read-only token for one resource instead of the user's session | More plumbing per feature |
| The untrusted input | Do not feed third-party content into a context that holds tools; summarize it in a separate, tool-less call first | Two model calls instead of one |
The third row is the one most teams can actually adopt. Split the work: one call reads the untrusted document and has no tools at all, and returns structured data. A second call has the tools but never sees the raw document — only the structured output of the first. An injection in the document can corrupt the summary, which is bad, but it cannot reach a tool, which is the difference between a wrong answer and an exfiltration.
How to check your own system
Draw the tool list. For each tool ask one question: can this reach a network the attacker can observe? An HTTP client obviously can. So does an image renderer that loads remote URLs, a Markdown renderer that produces clickable links, a logging sink the attacker can read, and an email tool. Then ask where the untrusted text enters. If the two sets meet inside one context window that also holds a privileged token, you have the trifecta.
The payload does not need to be clever to prove it. A benign marker is enough — a unique string you place in the "private" data, and a request that the agent append it to a URL. If the marker shows up in your own request log, the channel is real, and you have demonstrated the mechanism without touching anyone's data.
Where to go next
Run the free AI Audit to score your app against the OWASP LLM Top 10 and check whether the EU AI Act's transparency obligations reach you — it runs entirely in your browser, and nothing leaves your device.
If you want the full taxonomy — indirect injection, RAG poisoning, system-prompt extraction, insecure output handling, and the agent and MCP chapter this post is drawn from — that is Breaking AI, and the range is where you practice landing these against a target that is safe to break.