Prompt Injection via Database Content in Agentic Pipelines

Attackers exploit trusted database retrieval to inject commands agents follow automatically.

Senior Contributor · · 11 min read
Cover illustration for “Prompt Injection via Database Content in Agentic Pipelines”
Agent-Safe Data Access · October 9, 2026 · 11 min read · 2,438 words

Indirect prompt injection gets past the defenses most teams already built for direct injection, because it arrives through a trusted retrieval path that no one thought to watch. That distinction, not some new exotic technique, is what makes the database content threat worth taking seriously on its own terms.

The comparison to SQL injection helps explain why this matters so much, but it also shows where the analogy breaks down. SQL injection got solved structurally: parameterized queries separate code from data so that the two live in different layers and execute differently, no matter what a malicious string contains. Prompt injection has no equivalent fix sitting inside the model. A system prompt and a piece of user input arrive at the model as the same thing: natural-language text, parsed the same way, weighted the same way. There is no syntactic boundary inside a language model that marks one span of text as an instruction from the developer and another as data to be read but not obeyed. The defenses that exist today, pattern filters, classifiers, scoped permissions, all operate around the model, at the application and context layers. None of them change what happens inside the forward pass itself.

That absence of an internal boundary is also why the distinction between direct and indirect injection carries so much weight. If an attacker can reach the agent's input at all, that attacker is already an authenticated user, operating inside whatever identity and role-based access controls the system enforces. Even a successful direct injection is boxed in by permissions the attacker already had to live within. Indirect injection skips that requirement. The attacker never touches the agent directly. Instructions get planted inside content the agent will process on its own, a webpage, a document, an email, a database record, and the agent follows them because it has no way to tell those instructions apart from the legitimate material surrounding them.

The consequences scale with what kind of system is reading the poisoned content. In a chatbot, a successful injection produces a wrong answer, and the damage stays inside that one conversation. In an agentic system, the same injection produces a wrong action: a database query, a file write, an email sent, an API called. The only thing containing that action is whatever permissions the agent happens to be running with at the time. So the shift from bad output to bad action is why database-content injection needs its own treatment, not the same bucket as every other prompt injection concern.

How database records become injection payloads

Any record an agent pulls from a database is a potential payload, because the agent has no reason to treat retrieved content differently than it treats a direct instruction. Once that record lands in context, it reads like part of the task.

The path is mechanical and repeats the same way across systems. An agent receives a task, then queries a database or pulls chunks from a retrieval store as a normal part of doing its job. Whatever comes back, a customer support ticket, a product description, a chunk of a document, a note left in a CRM, is placed in the context window right next to the system prompt that's supposed to be steering the agent's behavior. Nothing in that context window marks the retrieved record as lower-trust than the developer's own instructions. Whatever natural-language text is embedded in that record gets processed and, if it reads as an instruction, followed.

Retrieval-augmented generation pipelines carry the highest concentration of this risk, because every chunk sitting in the retrieval corpus is a candidate payload. Agents that pull from RAG stores, browse the web, read email, or work inside code repositories are exposed in a way that agents confined to a single, controlled input are not. A poisoned chunk doesn't need to rank first in a retrieval call to do damage. It only needs to get pulled into context once.

A standardized protocol for connecting agents to external tools widens this surface. An agent reads and trusts by default tool descriptions, tool outputs, memory stores, and retrieval results, all as channels. A compromised or malicious MCP tool can hand back adversarial content as part of an ordinary tool response, and that response enters the model's context exactly the way a legitimate tool output would, no flags, no separate handling. Agents built with persistent memory add one more wrinkle: a single poisoned entry in that memory can affect every future session that draws on it, long after the original attack.

One documented case shows how little infrastructure you need to pull this off. Each agent had its own injection surface: the PR title reached Claude Code, an issue comment reached Gemini CLI, and an HTML comment buried in an issue body reached GitHub Copilot. All three agents read that content as part of their normal task context, followed the planted instructions, and exfiltrated GitHub Actions secrets, with Claude Code posting results back through a PR comment, Gemini CLI through an issue comment, and GitHub Copilot through a git commit. There is no command-and-control server, no external payload host, no separate delivery mechanism. The text already sitting inside the repository was the entire attack.

Agentic pipelines treat retrieved data as instructions by default

A language model is useful because it can process anything written in natural language, but that same property turns a retrieved database record into an unconditional instruction the moment it lands in context.

There's no syntactic boundary inside the model separating a developer's system prompt from a record a tool call just fetched. In conventional software, logic is written as code and user input arrives as a string, and the two live in separate layers that execute under different rules. A language model collapses that separation. A system prompt and a database record both show up as the same kind of artifact, natural-language text, and the model processes both the same way, with the same attention mechanism, the same weighting. A record pulled from a database by a tool call carries no less authority in that context window than an instruction the developer wrote by hand.

This isn't a hypothetical risk anymore. Unit 42's telemetry on web-based indirect prompt injection identified attacker intents spanning data destruction, denial of service, unauthorized transactions, leakage of sensitive information, and leakage of the system prompt itself. Across that telemetry, researchers identified 22 distinct techniques in active use for building and hiding payloads inside web content, the same category of content that database-connected agents pull from constantly when they browse, summarize, or retrieve. One of the cases in that telemetry marked the first reported detection of AI-based ad review evasion, a sign that attackers are actively finding new ways to abuse the retrieval path.

MCP agents cannot fix this just by routing retrieval through a tool call, because the tool itself can act as a confused deputy. MCP servers expose tools along with descriptions of those tools, and the model reads those descriptions as instructions about how and when to use them. The descriptions themselves are part of the attack surface before any data even gets retrieved. When an agent calls a tool and gets a response back, that response enters the context window as trusted tool output no matter what it actually contains. The agent is acting as a deputy, authorized to retrieve information on someone else's behalf, but the content it retrieves can carry instructions that the principal who authorized the retrieval never issued and never approved.

OWASP's Top 10 for Large Language Model Applications names prompt injection LLM01:2025, and it has topped the list for the second edition running. The organization states that no technique can guarantee complete mitigation, a consequence of the stochastic nature of language models themselves.

Write access and tool chaining amplify a successful injection

Once an injected instruction lands, the damage it can do is set entirely by what tools the agent is allowed to call next. Permission scope, not the cleverness of the injection, sets whether the result is a nuisance or a breach.

An agent has as many ways to cause harm as it has tools connected to it, and each tool is a door the attacker can walk through once an instruction takes hold. Data exfiltration can occur through channels that never appear in the agent's visible output at all: an image URL fetched automatically, a webhook call, an outbound API request, an email sent on blind copy. Unauthorized transactions are already happening in the wild. In May 2026, an attacker used obfuscated Morse code encoding to slip an injected instruction past safety filters built to catch plain natural-language attacks, and directed an agent to move funds after its transfer authorization had been unlocked through a gifted NFT. Privilege escalation follows the same logic in a simpler form: an agent that can read files and make HTTP requests represents a far larger exposure than one that can only return text to a user.

Write access to any channel the agent touches multiplies what a single successful injection can do. Message queues, databases, file systems, outbound network calls, any of these becomes a path for corruption or exfiltration once an agent can write to it. An event-streaming architecture, for instance, turns into an exfiltration surface the moment the language model driving it has write access to the stream itself.

The riskiest configuration combines two things that are each manageable on their own: an agent with access to privileged internal tools that also connects out to external MCP servers. An agent that can read production database records and also make outbound HTTP requests brings the injection surface, the retrieved content, and the exfiltration surface, the outbound call, into a single process with nothing standing between them. That combination is also what makes goal hijacking possible across multi-step pipelines: instead of triggering one bad action or pulling out one piece of information, an attacker can redirect the agent's entire objective across a chain of steps the agent executes on its own.

The EchoLeak attack against Microsoft 365 Copilot shows the full chain end to end. Hidden instructions sitting inside an email caused Copilot to query enterprise files with its own built-in search tool, then exfiltrate the results to a server the attacker controlled, by embedding the stolen data in an image URL the client loaded automatically. The user never typed a malicious prompt and never clicked a malicious link. The user never even had to open the content deliberately, because tools the user had already granted the agent permission to use carried out the whole attack.

Application-layer defenses do not close the database content threat

Defenses built to screen user input at the application boundary miss this entire category of attack, because by the time a retrieved database record reaches the model, it has already passed through every filter the system has, flagged as legitimate data.

Standard input-layer defenses are built for a different channel than the one this threat uses. Regex pattern matching, calls out to moderation APIs, semantic classifiers run against incoming prompts, all of these are calibrated to catch direct injection: attacker-controlled text arriving at the model's front door, where a user is typing. A database record pulled mid-session by a tool call never passes through that front door. It enters the context window as tool output, after the input layer has already cleared the session and moved on. The filter built to catch malicious prompts never gets a chance to see it.

Obfuscation makes this worse. Encoding schemes built to slip past filters trained on plain natural-language attack patterns keep evolving, and content retrieved from a database or a web page can arrive in formats those filters were never trained to recognize.

The gaps compound as pipelines grow more complex. If a new microservice, agent, or model integration joins a pipeline, it has to reimplement the same checks on its own, and if one internal tool ships without them, that one surface stays wide open. Multi-agent pipelines make this worse still, because no one has a unified view across the whole pipeline of which piece of retrieved content triggered which violation. Unit 42's research found 22 distinct payload-delivery techniques already active in the wild for indirect prompt injection, including obfuscation methods that existing signature-based defenses don't cover, so the attack surface is changing faster than detection systems can keep pace with.

Instructing the system prompt to "ignore adversarial content in retrieved data" doesn't hold up reliably either. OWASP's own acknowledgment, that the stochastic nature of language models means no technique can guarantee complete mitigation, applies directly here. An instruction sitting in the system prompt is competing for the model's attention against an injected instruction sitting in retrieved content, and the model has no architectural mechanism that lets it weight one over the other with any guarantee. Detection-based defenses deployed at the input boundary are not useless, but they are structurally unable to intercept an injection that arrives through the content path instead of the input path, which is exactly the path database records travel.

Structural defense means never exposing raw production content to agents

You stop this threat structurally by making sure agents never retrieve raw, unmodeled production records. A governed, pre-calculated dataset has no freeform text fields left for an attacker to plant instructions inside, which removes the payload surface rather than trying to catch the payload after the fact.

The underlying principle is to treat everything an agent retrieves as data and never as instructions, and to enforce that distinction architecturally rather than hoping a system prompt will hold the line. Content coming from different trust levels should carry a tag that identifies its source. Retrieved documents, tool outputs, and messages passed between agents all need to be processed under different trust assumptions than instructions written directly by a developer. Prompting alone can't enforce any of this, because the model still has no internal boundary to lean on even when it's told one exists.

The structural answer models the data before an agent ever touches it. A governed, pre-calculated dataset, metrics computed ahead of time, aggregated, and served as structured Parquet, has no freeform prose fields left in it for an injection payload to hide inside. An agent asking for ARR in Q3 against a modeled dataset gets back a number and a definition of what that number means, not a customer-generated text record that could contain absolutely anything, including instructions written to look like data. That's the difference that matters: the attack surface isn't being detected and blocked after the fact, it's removed by construction, before the agent ever has a chance to read it.

Sources

  1. Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild
  2. Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
  3. Agent Data Injection Attacks are Realistic Threats to AI Agents
  4. Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
  5. Model Context Protocol (MCP): Security Design ...

More in Agent-Safe Data Access