A chatbot that says something wrong is embarrassing. An AI agent that does something wrong is an incident. The difference is that an agent holds credentials, reads content it did not write, and takes actions in real systems: sending email, updating records, moving files, calling APIs.
That combination makes AI agent security a problem traditional application controls were not designed for. This guide explains why prompt injection cannot simply be filtered out, what two real incidents teach, the three conditions that make data theft possible, and the controls that keep the damage small when the model is eventually fooled.
Why Agents Are a Different Security Problem
In ordinary software, code and data are kept apart. A database query cannot rewrite the program that issued it. A language model has no such separation: its instructions and the content it reads arrive as the same stream of text. Any email, document, web page or ticket the agent reads can contain text that looks like an instruction, and the model may follow it. This is prompt injection.
The OWASP Top 10 for LLM Applications ranks prompt injection first (LLM01), and lists Excessive Agency (LLM06) — giving a model more permissions, functions or autonomy than its task needs — among the most critical risks. In December 2025 OWASP published a separate Top 10 for Agentic Applications, built by more than 100 security practitioners, covering systems that plan, keep memory, call tools and act with delegated authority.
Assume the model will occasionally be talked into something. Security comes from limiting what it can do when that happens, not from hoping it never does.
Two Incidents Worth Knowing
Neither incident needed a sophisticated attacker breaking into a server. In the first, the attack was an ordinary email. In the second, there was no attacker at all: an agent had permissions it should never have had, in an environment with no barrier between testing and production.
Three Conditions That Make Theft Possible
The clearest way to reason about agent data leaks is what Simon Willison named the lethal trifecta in June 2025. An agent is open to data theft when it has all three of the following at once.
| Condition | Example | How to remove it |
|---|---|---|
| Access to private data | Reads the CRM, the shared drive, the inbox | Give the workflow only the records it needs, read through a narrow interface |
| Exposure to untrusted content | Processes inbound email, supplier documents, web pages | Separate the step that reads untrusted content from the step that touches private data |
| A way to send data out | Can send email, call arbitrary URLs, render remote images | Allow only fixed destinations, and strip links and remote images from output |
The design rule is simple: break at least one leg in every workflow. An agent that summarises inbound supplier email may need untrusted content and some private data, so it must not be able to contact arbitrary addresses. An agent that can email customers should not also be reading unvetted documents in the same context. EchoLeak worked because all three legs were present together.
Controls That Actually Limit Damage
What Filters Cannot Do
Prompt-injection detectors and guardrail models are worth having. They catch obvious attacks and raise the cost of subtle ones. But they are probabilistic: they score text and make a judgement, and a determined attacker can phrase an instruction they miss. EchoLeak got past a production classifier built by one of the best-resourced security teams in the world.
That is why the controls above are architectural rather than linguistic. A filter tries to stop the model being fooled. Permissions, allowlists, approvals and environment separation decide what a fooled model can actually do. The second approach still works on the day the first one fails.
The same logic applies wherever the model runs. A self-hosted model removes the third-party provider from the picture, but it is just as susceptible to instructions hidden in a document. Data location and agent security are separate questions.
Five Questions Before Go-Live
What can it read?
List every data source the agent can reach, including through search or retrieval. If the honest answer is “most of the shared drive”, the scope is too wide.
What untrusted content reaches it?
Anything written by someone outside the organisation: email, attachments, form submissions, web pages, supplier documents. Each one is a way to give the agent instructions.
What can it change?
Every write, send, delete and permission change it can make, and which of those are irreversible. Irreversible actions need an approval step.
Where can it send data?
Email, webhooks, URLs, rendered images, files written to shared locations. If questions one, two and four all have broad answers, the trifecta is complete.
How would we know, and how would we undo it?
Which alert fires, who receives it, what the logs show, and how the change is rolled back. If nobody can answer, the workflow is not ready.
Security of this kind is not new to connected systems. The same principles of least privilege, device identity and audit run through our work on IoT security for connected products; agents simply add a component that can be persuaded.
Frequently Asked Questions
Conclusion
AI agent security is less about clever detection and more about old-fashioned engineering discipline: least privilege, narrow interfaces, separate environments, approval for what cannot be undone, and a complete record of what happened. The model will be fooled eventually. The job is to make sure that, when it is, it cannot reach far.
Agents with bounded authority, audit trails and rollback are how we build automation, and the reason choosing between automation, RPA and agents is partly a security decision. If you are planning to give an agent real permissions, our AI automation services start with the threat model.
