Prompt injection explained: How attackers hijack your AI tools without touching your network
As AI assistants become part of everyday work, we are asking them to read more on our behalf. An email thread that would take twenty minutes to catch up on becomes a short summary. A contract gets compared against an internal brief in seconds. Some assistants go further, drafting a reply or updating a record in another system once they have finished the analysis.
These are genuinely useful shortcuts. They also raise a question that is easy to overlook: what happens if something the assistant reads contains instructions from someone else?
Picture this. You ask an assistant to summarize a document a client sent over. Alongside the legitimate content, the document contains a note claiming the task requires an external verification step and instructs the assistant to upload your internal files to another website. You asked for a summary. The document is trying to decide what happens to your data.
That is a hypothetical example of indirect prompt injection. The attacker places instructions inside material an AI system will process, hoping the assistant treats them as though they came from an authorized source. The attempt may not succeed, and the consequences depend on what the assistant can access and what safeguards are in place. But the attacker may not need to steal a password or compromise your network first. Getting the right content in front of the assistant can be enough.
How the instructions get mixed up
An AI assistant works with several kinds of input at once: instructions from its developer, your request, and any material it retrieves to complete that request. That material might be a webpage, a document, or results pulled from a search across company files.
The application can label these sources and tell the model which instructions take priority. The difficulty is ensuring the model consistently respects those boundaries when the retrieved material itself contains language that looks like a command. A paragraph in a document should remain something to analyze, even when it claims to be an urgent instruction from an administrator.
This is where prompt injection differs from an ordinary misdirected request. In a direct injection, someone submits instructions to the AI application itself to try to change its behaviour. In an indirect injection, the instructions arrive through content the assistant is reading while carrying out a completely legitimate task. The employee may have done nothing unusual beyond asking for a summary.
The wording does not have to be an obvious “ignore previous instructions.” It might claim a particular action is necessary to finish the work, giving the assistant a plausible-sounding reason to depart from what you actually asked for. Whether the assistant follows that instruction depends on both the model and the controls around it.
What researchers have demonstrated
In May 2023, security researcher Johann Rehberger published two related proofs of concept against ChatGPT plugins. In the first, he showed that malicious instructions embedded in a webpage, retrieved via the WebPilot browsing plugin, could trigger the Zapier plugin to access a user’s Gmail, summarize recent emails, and send that data to an attacker-controlled URL. The user saw nothing unusual.
In the second variant, he demonstrated that an AI assistant rendering markdown could be made to construct an image URL containing conversation data, sending it to an external server the moment the interface loaded the image. Neither attack required the user to click a suspicious link or open a malicious file.
These were proofs of concept tied to specific plugin integrations available at the time, not descriptions of how every AI tool works today. But they illustrated the core mechanism clearly, and the underlying vulnerability has not gone away.
Researchers from ETH Zurich and Invariant Labs examined the broader problem through AgentDojo, a benchmark introduced at NeurIPS 2024. They tested AI agents performing tasks involving email, calendars, and other tools, with attacks embedded in the information those tools returned. The evaluations included attempts to make agents leak emails. These were controlled experiments, not reports of a customer breach, but they gave researchers a structured way to study what happens when useful information and malicious instructions arrive together.
Why connected tools raise the stakes
An assistant that only produces text can still cause problems if an injection makes it disclose sensitive context or return a misleading answer. A manipulated document summary might miss an unfavorable clause. A tampered recommendation might favor one option without any obvious reason. No external data transfer is required for the output to become unreliable.
The consequences become broader when the assistant has tools available. An agent with permission to read documents and send messages may be able to turn a malicious instruction into an unwanted action. A successful injection does not automatically give an attacker every permission the assistant holds. The outcome still depends on which tools are available, what they allow, and whether another control blocks the request before it executes.
For anyone approving an AI deployment, those specifics matter more than the label “agentic.” An assistant that drafts a message for a human to review presents a different risk profile from one that can choose a recipient and send it independently. The difference starts with checking the actual configuration, not the marketing description.
Where an assistant might encounter an injection
Documents are the most obvious route. A shared PDF, a contract, a spreadsheet sent by a third party may contain instructions alongside the content your employee wants to analyze. Moving a file into a company folder does not make everything inside it trustworthy.
Email is another: An assistant working through your inbox is reading material written entirely by external parties. A message can attempt to influence how the assistant handles other messages, what it includes in a reply, or what actions it takes next.
Customer records and support tickets present the same exposure if an assistant retrieves text that someone outside your organization can edit or submit.
Web browsing adds a further surface. Malicious instructions may appear in visible page text, in HTML that is less apparent to the human eye, or in content that the application’s retrieval method surfaces to the model even though it is not visible on screen. What an employee sees and what the assistant processes are not always the same thing.
Visual inspection is not a reliable defense against any of these. Employees cannot be expected to identify every malicious instruction in every piece of content an assistant processes. The controls need to work even when the suspicious content goes unnoticed.
What organizations can do
Start with access. Check what each assistant can actually read, write, and change. Remove connections it does not need, and grant only the permissions required for its intended tasks. Where reading is sufficient, avoid granting write access simply because the integration offers it.
Make approvals specific. A prompt asking “Approve the next step?” gives the reviewer almost nothing to work with. Before an assistant shares a file externally, the person approving should see the destination and the content being sent. A destination supplied inside a document should not quietly become an authorized destination. Enforce approval requirements at the application level, not just through instructions telling the model to ask permission.
Isolate retrieved content from trusted instructions. Keep external material separate in the workflow, and limit its ability to influence tool use. Where possible, separate the part of the workflow that reads external content from the capabilities that can make sensitive changes. Validate tool requests and restrict outbound destinations.
Test with adversarial content. Include malicious documents and emails in a controlled test environment, and check whether the application stops an unwanted action even when the model requests it. Keep records of tool calls and outcomes so unexpected behavior can be investigated. Repeat relevant tests when permissions, integrations, or underlying models change.
Ask your providers the right questions. Find out how they support these controls and what administrators can inspect. Prompt injection remains an active area of research. Model improvements reduce the risk but do not remove the need for safeguards around connected tools.
What employees should watch for
If an assistant requests an unexpected upload, proposes a new recipient, or asks for permissions that were not part of the original task, pause before approving. Check whether the proposed action actually follows from what you asked for, and look at the information it is proposing to share.
For important summaries or recommendations, check the output against the source material rather than taking the assistant’s confidence at face value. If something looks wrong, stop the workflow and report it through your organization’s security channel. Preserve the conversation and the relevant document or link, following your team’s instructions for handling sensitive evidence. You do not need to prove an injection occurred before raising a concern.
Awareness helps people notice when something is off. Permissions, approval checks, and monitoring protect them when there is no obvious warning. As assistants take on more work with company information, both are necessary.
Sources
[1] OWASP LLM01 2025 Prompt Injection
[3] Debenedetti et al., AgentDojo, 2024
[4] Anthropic, Mitigating the risk of prompt injections in browser use, 24 November 2025
[5] Microsoft, Defend against indirect prompt injection attacks, 24 March 2026