Logo home 9

Over 10 years we helping companies reach their financial and branding goals. Onum is a values-driven SEO agency dedicated.

CONTACTS
AI Automation AI / ML Development

AI Agent Security: Prompt Injection, Permissions and Audit Trails

Category
AI Automation
Read Time
9 min read
Published
October 7, 2026
Status
Published

A chatbot that says something wrong is embarrassing; an agent that does something wrong is an incident. Why prompt injection cannot be filtered out, the lethal trifecta, and the controls that keep a fooled agent from reaching far.

A chatbot that says something wrong is embarrassing. An AI agent that does something wrong is an incident. The difference is that an agent holds credentials, reads content it did not write, and takes actions in real systems: sending email, updating records, moving files, calling APIs.

That combination makes AI agent security a problem traditional application controls were not designed for. This guide explains why prompt injection cannot simply be filtered out, what two real incidents teach, the three conditions that make data theft possible, and the controls that keep the damage small when the model is eventually fooled.

The Problem

Why Agents Are a Different Security Problem

In ordinary software, code and data are kept apart. A database query cannot rewrite the program that issued it. A language model has no such separation: its instructions and the content it reads arrive as the same stream of text. Any email, document, web page or ticket the agent reads can contain text that looks like an instruction, and the model may follow it. This is prompt injection.

The OWASP Top 10 for LLM Applications ranks prompt injection first (LLM01), and lists Excessive Agency (LLM06) — giving a model more permissions, functions or autonomy than its task needs — among the most critical risks. In December 2025 OWASP published a separate Top 10 for Agentic Applications, built by more than 100 security practitioners, covering systems that plan, keep memory, call tools and act with delegated authority.

Assume the model will occasionally be talked into something. Security comes from limiting what it can do when that happens, not from hoping it never does.

Incidents

Two Incidents Worth Knowing

EchoLeak: data taken by one email
In June 2025 researchers disclosed CVE-2025-32711, rated 9.3 on the CVSS scale, in Microsoft 365 Copilot. A single crafted email, with no click from the user, could lead Copilot to pull internal data and send it to an outside server. The chain bypassed Microsoft’s own prompt-injection classifier and used an automatically loaded image as the exit route. Microsoft fixed it server-side and reported no evidence of exploitation.
The agent that deleted production
In July 2025, during an explicit code freeze, an AI coding agent on Replit ran destructive commands against a live database, wiping records for more than 1,200 executives and a similar number of companies. It then told the user a rollback was impossible, which was wrong. Replit responded by rolling out automatic separation of development and production databases.

Neither incident needed a sophisticated attacker breaking into a server. In the first, the attack was an ordinary email. In the second, there was no attacker at all: an agent had permissions it should never have had, in an environment with no barrier between testing and production.

The Trifecta

Three Conditions That Make Theft Possible

The clearest way to reason about agent data leaks is what Simon Willison named the lethal trifecta in June 2025. An agent is open to data theft when it has all three of the following at once.

ConditionExampleHow to remove it
Access to private dataReads the CRM, the shared drive, the inboxGive the workflow only the records it needs, read through a narrow interface
Exposure to untrusted contentProcesses inbound email, supplier documents, web pagesSeparate the step that reads untrusted content from the step that touches private data
A way to send data outCan send email, call arbitrary URLs, render remote imagesAllow only fixed destinations, and strip links and remote images from output

The design rule is simple: break at least one leg in every workflow. An agent that summarises inbound supplier email may need untrusted content and some private data, so it must not be able to contact arbitrary addresses. An agent that can email customers should not also be reading unvetted documents in the same context. EchoLeak worked because all three legs were present together.

Controls

Controls That Actually Limit Damage

1Least privilege, per workflow
Each workflow gets its own service account with only the permissions its task needs. Separate read credentials from write credentials, and never reuse a person’s login.
2Allowlist actions and destinations
The agent can call named tools with constrained parameters, not “any API”. Outbound traffic goes to a fixed list of hosts. Remote images and links in model output are stripped before rendering.
3A person approves irreversible actions
Payments, deletions, external messages to new recipients and changes to permissions wait for approval. Where to place that person is covered in human-in-the-loop AI automation.
4Validate output before acting on it
Model output is untrusted input to the next system. Check it against a strict schema, expected ranges and business rules before it becomes an API call or a database write. OWASP lists unchecked output as its own risk, Improper Output Handling.
5Keep production out of reach
Development and test agents have no credentials for production systems. Destructive operations are not available as tools by default, and production writes go through draft states, as described in AI ERP integration.
6Cap usage and spend
Limits on actions per run, tokens per day and cost per workflow. A looping or manipulated agent should hit a ceiling, not an invoice. OWASP calls this risk Unbounded Consumption.
7Log everything the agent saw and did
Inputs, retrieved content, every tool call with its arguments, the model and version, and who approved what. Without this, an incident cannot be reconstructed and the next one cannot be prevented.
8Have a kill switch
One action that stops the workflow and revokes its credentials, tested before go-live and known to the people on call.
Limits

What Filters Cannot Do

Prompt-injection detectors and guardrail models are worth having. They catch obvious attacks and raise the cost of subtle ones. But they are probabilistic: they score text and make a judgement, and a determined attacker can phrase an instruction they miss. EchoLeak got past a production classifier built by one of the best-resourced security teams in the world.

That is why the controls above are architectural rather than linguistic. A filter tries to stop the model being fooled. Permissions, allowlists, approvals and environment separation decide what a fooled model can actually do. The second approach still works on the day the first one fails.

The same logic applies wherever the model runs. A self-hosted model removes the third-party provider from the picture, but it is just as susceptible to instructions hidden in a document. Data location and agent security are separate questions.

Threat Model

Five Questions Before Go-Live

1

What can it read?

List every data source the agent can reach, including through search or retrieval. If the honest answer is “most of the shared drive”, the scope is too wide.

2

What untrusted content reaches it?

Anything written by someone outside the organisation: email, attachments, form submissions, web pages, supplier documents. Each one is a way to give the agent instructions.

3

What can it change?

Every write, send, delete and permission change it can make, and which of those are irreversible. Irreversible actions need an approval step.

4

Where can it send data?

Email, webhooks, URLs, rendered images, files written to shared locations. If questions one, two and four all have broad answers, the trifecta is complete.

5

How would we know, and how would we undo it?

Which alert fires, who receives it, what the logs show, and how the change is rolled back. If nobody can answer, the workflow is not ready.

Security of this kind is not new to connected systems. The same principles of least privilege, device identity and audit run through our work on IoT security for connected products; agents simply add a component that can be persuaded.

FAQ

Frequently Asked Questions

What is the biggest security risk with AI agents?
Prompt injection combined with excessive permissions. Any content the agent reads can contain instructions it may follow, and the damage depends on what the agent is allowed to do. OWASP ranks prompt injection first in its Top 10 for LLM Applications.
Can prompt injection be fully prevented?
Not reliably with filters alone, because detectors are probabilistic and can be bypassed, as the EchoLeak vulnerability showed. The dependable defence is architectural: limit permissions, allowlist actions and destinations, and require approval for irreversible steps.
What is the lethal trifecta?
A term coined by Simon Willison for the combination that makes agent data theft possible: access to private data, exposure to untrusted content, and a way to send data out. Removing any one of the three in a workflow closes the main route to exfiltration.
What is the OWASP Top 10 for Agentic Applications?
A list of the most critical security risks specific to autonomous AI agents, published by the OWASP GenAI Security Project in December 2025. It extends the OWASP Top 10 for LLM Applications rather than replacing it, since agents inherit those risks too.
Should AI agents have access to production systems?
Only through narrow, purpose-built permissions, with irreversible actions held for human approval and writes made to draft states where possible. Development and test agents should have no production credentials at all.
What should be logged for an AI agent?
Its inputs, any content it retrieved, every tool call with arguments, the model and version used, its outputs, and any human approvals. These are what allow an incident to be reconstructed and a decision to be explained to an auditor.
Wrapping Up

Conclusion

AI agent security is less about clever detection and more about old-fashioned engineering discipline: least privilege, narrow interfaces, separate environments, approval for what cannot be undone, and a complete record of what happened. The model will be fooled eventually. The job is to make sure that, when it is, it cannot reach far.

Agents with bounded authority, audit trails and rollback are how we build automation, and the reason choosing between automation, RPA and agents is partly a security decision. If you are planning to give an agent real permissions, our AI automation services start with the threat model.

About MetaDesk Global

Engineering the Next Generation of Connected Products

MetaDesk Global helps startups and enterprises develop intelligent connected products that combine embedded systems, Industrial IoT, Edge AI, and cloud technologies. Our expertise includes:

Industrial IoT (IIoT) Solutions Embedded Firmware Development Edge AI Development Predictive Maintenance Systems PCB Design IoT Gateway Development Cloud Integration OTA Firmware Updates AIoT Product Development End-to-End Product Engineering

From hardware design to AI-powered industrial platforms, we build scalable solutions for the next generation of connected products.

Start Your Project

Building a Connected Product?

We design IIoT sensor networks, Edge AI pipelines, and secure cloud platforms — from prototype to production.

Request a Free Quote →