← All writing
LLM security

The lethal trifecta — why your agent is one email away from exfiltration

An LLM agent becomes dangerous the moment it holds three things at once: private data, exposure to untrusted content, and a way to talk to the outside world. Here is why that combination leaks, and what to actually remove.

A chatbot that can only talk is a low-stakes system. Give it your inbox, a browser, and an HTTP tool, and you have built something that can be instructed by a stranger and will act on those instructions with your credentials. Nothing was hacked. The model did exactly what it was designed to do.

The failure has a shape, and once you can see the shape you can find it in almost any agent architecture in about ten minutes.

The three ingredients

An agent is exposed when it simultaneously has:

  1. Access to private data — your email, your repos, your customer records, your internal wiki. Anything the attacker cannot read directly.
  2. Exposure to untrusted content — a web page it fetches, an email it summarizes, a PDF a user uploads, an issue comment, a calendar invite. Any text that arrived from someone who is not the operator.
  3. A way to communicate externally — an HTTP tool, an email send action, a webhook, a rendered image with a remote URL, even a Markdown link the user might click.

Hold all three and the attack writes itself: untrusted content carries instructions, the model follows them, private data leaves through the exit.

The model has no reliable way to tell your instructions from the data it is reading. They arrive in the same context window, in the same language, with the same authority. "Ignore the above" is not a bug in the prompt — it is a consequence of the architecture.

Why guardrails do not close it

The instinct is to filter. Scan the incoming content for injection attempts, add a system prompt that says never follow instructions found in documents, maybe run a classifier over the output.

These help at the margin and none of them are a boundary. The reason is counting: a filter has to catch every phrasing in every language and encoding, forever, while the attacker only needs one that works. Base64 the payload. Put it in alt text. Split it across two documents the agent will read in the same session. Write it as a plausible-looking system notice. The asymmetry never moves in the defender's favour.

Worse, a filter that mostly works is more dangerous than no filter, because it converts a known-open door into one everybody believes is shut.

Remove an ingredient instead

The trifecta is a conjunction, so you break it by deleting a term, not by inspecting one harder.

RemoveHow it looks in practiceWhat it costs
The exitNo outbound HTTP from the agent's tool set; allowlist the two domains it genuinely needsThe agent cannot "just fetch" things
The private dataGive the agent a scoped, read-only token for one resource instead of the user's sessionMore plumbing per feature
The untrusted inputDo not feed third-party content into a context that holds tools; summarize it in a separate, tool-less call firstTwo model calls instead of one

The third row is the one most teams can actually adopt. Split the work: one call reads the untrusted document and has no tools at all, and returns structured data. A second call has the tools but never sees the raw document — only the structured output of the first. An injection in the document can corrupt the summary, which is bad, but it cannot reach a tool, which is the difference between a wrong answer and an exfiltration.

How to check your own system

Draw the tool list. For each tool ask one question: can this reach a network the attacker can observe? An HTTP client obviously can. So does an image renderer that loads remote URLs, a Markdown renderer that produces clickable links, a logging sink the attacker can read, and an email tool. Then ask where the untrusted text enters. If the two sets meet inside one context window that also holds a privileged token, you have the trifecta.

The payload does not need to be clever to prove it. A benign marker is enough — a unique string you place in the "private" data, and a request that the agent append it to a URL. If the marker shows up in your own request log, the channel is real, and you have demonstrated the mechanism without touching anyone's data.

Where to go next

Run the free AI Audit to score your app against the OWASP LLM Top 10 and check whether the EU AI Act's transparency obligations reach you — it runs entirely in your browser, and nothing leaves your device.

If you want the full taxonomy — indirect injection, RAG poisoning, system-prompt extraction, insecure output handling, and the agent and MCP chapter this post is drawn from — that is Breaking AI, and the range is where you practice landing these against a target that is safe to break.

LLM securityprompt injectionagents