Prompt Injection Risks in Multi-Model Writing Pipelines
Multi-agent writing systems inherit vulnerabilities from passing unverified outputs between agents.

Prompt injection in a multi-model writing pipeline, run through an entire chain of agents rather than a single checkpoint, is a different failure mode from prompt injection in a single model scaled up, since adversarial content that clears one checkpoint can travel through the whole chain as though it belonged there. That distinction, and what it means for teams building automated writing systems, is the subject of this piece.
Multi-model pipelines face a structurally different injection problem than single-model deployments
A single-model deployment has one clear failure boundary. A model receives input, produces output, and whatever goes wrong is contained inside that one exchange. If an injected instruction slips past the model's guardrails, the damage appears in a single response, and whoever reviews that response can catch it.
A multi-agent pipeline removes that boundary. Once adversarial content clears the first agent in the chain, every agent after it treats that output as legitimate input, not as something to be re-examined. Research on multi-agent LLM pipelines by a team including Faisal Haque Bappy and coauthors describes this as a security gap absent in single-agent settings: a pipeline that passes intermediate outputs between specialized agents creates a structure where adversarial content, once accepted anywhere in the chain, is propagated as trusted input throughout the rest of it.
The distinction matters because it changes where the fix has to live. Swapping in a more capable model at any stage of the pipeline does nothing to close this gap, because the vulnerability doesn't come from any one model's reasoning limits. It comes from the architecture itself: agents built to exchange outputs without verifying them. A writing pipeline built from a summarizer agent, a critique agent, and a drafting agent, each one handing its output to the next, has this exact structure baked in. Each stage trusts the one before it by design, and that design is what an attacker exploits.
The three unverified boundaries where adversarial content enters a multi-agent writing pipeline
The Bappy et al. research maps this exposure onto three categories of inter-agent boundaries, and each one is left unverified in current architectures.
The first is the content boundary. Retrieved or ingested external content can override an agent's directives with no sanitization step in between, because the line between data and executable instruction simply isn't enforced at the point where content enters an agent's context.
The second is the identity boundary. A downstream agent in most pipelines infers who sent it a given piece of content based on where in the routing sequence that content arrived, not on any credential it can check. That means a downstream agent has no way to confirm that what it's receiving actually came from the legitimate upstream agent, rather than from injected content made to look like it did.
The third is the delegation boundary. Downstream agents aren't required to stick to a prescribed execution plan, so an injection that alters one agent's output partway through the chain can quietly redirect the entire workflow without anything flagging the deviation.
What ties these three together is the absence of one shared security primitive: boundary verification. When data crosses from one agent to the next, you need explicit checks on content, identity, execution intent, and state integrity. No current writing pipeline architecture enforces this by default, which is precisely what research from UCL and Stanford on Prompt Infection demonstrates in practice: malicious prompts self-replicate across interconnected agents and spread silently, even in pipelines where the agents involved don't publicly share every communication with each other.
Self-replicating prompt injection moving through a pipeline without triggering any existing safeguard
If a prompt injection lands anywhere in a multi-agent writing pipeline, it can copy its own instructions into every agent's context downstream of it. The Prompt Infection research gives this behavior a name and treats it with the seriousness that behavior deserves: adversarial prompts replicate across interconnected LLMs the way a virus replicates across hosts, and the threats that follow include data theft, scams, misinformation, and disruption across the whole system, all without any single point of human intervention triggering along the way.
The mechanics are straightforward once laid out. Once an injected document reaches Agent A, the summary it produces carries those instructions forward inside it. Agent B receives that summary, treats it as ground truth because that's what summaries from upstream agents are supposed to be, and acts on the embedded directives now sitting inside it. Agent C then executes based on Agent B's already-corrupted output. No agent in this chain flags an anomaly, because each one is receiving what a valid upstream output is supposed to look like.
Existing safeguards miss this because of where they sit. Input sanitization checks what a user types into the first field of the pipeline. It has no visibility into what Agent A hands to Agent B, because that handoff never passes through a user-facing boundary. The Prompt Infection researchers confirm that multi-agent systems stay highly exposed to this even when the agents involved don't share all their communications publicly with each other. If you keep agent communication private, that alone doesn't make a pipeline safer.
For a writing pipeline specifically, a critique agent is a high-value target for this kind of propagation. A critique agent that produces detailed feedback on a draft is, by the nature of its job, built to absorb content from a source document and restate it in a new form. Injected instructions sitting in that source document can ride straight through the critique agent's output, so they land directly in a final draft agent's context window, fully dressed up as legitimate editorial feedback.
Brand Corpora and Ingested Documents as the Highest-Risk Injection Surface
For content teams running RAG-based writing pipelines, the indirect injection vector through ingested documents is more dangerous than injection through direct user input, because it comes down to which defenses teams actually have in place. Direct injection arrives through the field a user types into, and most teams at least check that field. Indirect injection arrives buried inside a retrieved document, so most teams validate the query a user submits but trust everything already sitting in the knowledge base without question.
The architecture itself makes this worse. In a RAG setup, retrieved documents get concatenated directly into the model's context window right alongside the user's legitimate query, and no additional separation is enforced between instruction and data at that stage. Whatever text a retrieved document contains is placed in the same context as the query itself, and the model has no structural reason to treat the two differently.
For a writing pipeline, that makes the corpus itself the attack surface. Tone-of-voice guides, past campaign copy, brand wikis, product briefs: any document an attacker can insert into that corpus, or quietly edit once it's already there, becomes a standing injection vector. A wiki page a contractor edited months ago, or a support transcript ingested automatically into the knowledge base, can sit there carrying an injected instruction indefinitely.
That persistence is what makes this vector so much more dangerous than a one-off injected query. A poisoned document doesn't get flushed out by a session reset. It stays in the corpus and gets retrieved again on every query that touches the same topic, firing the same injected instruction over and over rather than once. And because the query itself still looks clean and the retrieved content still looks like ordinary brand material, nothing about the exchange looks anomalous to anyone reviewing it after the fact.
The additional attack surface that opens when a writing pipeline accepts image inputs
Writing pipelines that accept images, whether product screenshots, brand assets, or design mockups, open an injection surface that no amount of text-layer sanitization can see. Vision encoders process an image holistically rather than parsing it the way a text sanitizer parses characters, so an instruction that would get flagged and blocked instantly as plain text can pass straight through undetected once it's encoded into pixels instead.
The Cloud Security Alliance's AI Safety Initiative identifies four distinct techniques for embedding instructions this way: typographic text, where the instruction is rendered as visible text inside the image itself; steganographic encoding, where the instruction is hidden inside pixel values invisible to the eye; adversarial pixel perturbations, imperceptible noise patterns that redirect model behavior; and physical-world signage [1][2][3]. Typographic injection is the simplest to picture: an instruction rendered in plain visible text on what otherwise looks like an ordinary product photo or screenshot, invisible to a reviewer skimming the image but fully legible to a vision-capable agent processing it. Each of the four techniques has its own stealth profile and needs its own defense, so no single filter on image inputs catches all four at once.
OWASP's 2025 revision of its LLM Top 10 has already caught up to this, extending its prompt injection classification (LLM01) explicitly to cover multimodal injection vectors, not just text. And the propagation risk from earlier sections applies here without modification: a malicious image processed by a vision-capable agent injects instructions that then travel downstream through the exact same inter-agent trust chain already described. Multimodal input is simply another entry point into pipeline propagation, not a separate problem from it.
The Cloud Security Alliance's recommendation is correspondingly simple to state and not yet widely followed: treat image inputs from untrusted sources with the same skepticism applied to user-supplied text. Most pipeline architects have built that skepticism into their text-handling and have not yet extended it to images.
What the EU AI Act's August 2026 watermarking requirement means for pipelines that pass content through multiple models
Multi-model writing pipelines carry a provenance problem, and as of this year, that problem is a matter of regulatory compliance, not engineering preference. The EU AI Act's watermarking provision became enforceable on August 2, 2026, with a grandfathering period running until December 2, 2026 for systems already on the market at that point. The provision requires AI model providers to watermark their output. If content marketing teams run multi-model pipelines commercially in the EU, they face a separate set of deployer disclosure obligations under Article 50(4), distinct from the provider watermarking requirement itself.
Anthropic announced in August 2026 that Claude would watermark output from any new model released on or after August 2, 2026, with older models scheduled to follow by December 2026. Anthropic also released a detection API in private preview that same month, made available to regulators, law enforcement, media organizations, fact-checkers, independent researchers, educational organizations, and EU civil society groups.
None of this solves the problem of attribution and propagation across the pipeline. A watermark applied to Agent A's output has no guarantee of surviving Agent B's rewrite and Agent C's final edit intact. Watermarks from intermediate models can be stripped or corrupted by each subsequent model pass in the chain, and a pipeline built from several models in sequence ends up with no reliable way to attribute its final output to any one specific model at all, which is exactly the outcome the regulation is meant to prevent.
This compounds the injection risk already described in earlier sections. An attacker who corrupts an intermediate agent's output in a propagation attack corrupts the watermark chain at the same time. Forensically tracing where in the pipeline the content was altered becomes effectively impossible once the chain has been broken.
How standard single-layer defenses fail at pipeline scale
The four conventional defenses against prompt injection, input sanitization, output filtering, prompt engineering, and model fine-tuning, each address one single point in a pipeline. None of them was built with inter-agent boundaries in mind, because none of them was designed for a pipeline. Research on multi-agent defense architectures confirms that these conventional approaches run into real limits handling novel attack vectors while still preserving system usefulness, a limitation that traces back to the fact that they were built for single-endpoint deployments.
Input sanitization only checks what enters the first agent in the chain. It has no visibility whatsoever into what Agent A sends to Agent B afterward, because that handoff happens entirely downstream of the only checkpoint input sanitization actually watches. Output filtering can catch obvious violations sitting in a pipeline's final output, but it's blind to subtler content-level injections, the kind that simply rewrite prose in a way that never trips a known attack signature.
Schema-level enforcement, restricting which tools an agent is even allowed to call, is the strongest structural control available today, but it constrains tool calls, and nothing else. In a writing pipeline, the output the pipeline produces is natural-language text, not a tool call, so schema restriction offers no protection at all against a content-level injection that never touches a tool.
The defense posture that's emerged by 2026 has moved to defense-in-depth: multiple layers stacked together, with the exact composition varying by framework but commonly including input sanitization, output filtering, prompt and context isolation, least-privilege tool access, and runtime monitoring, all running at once. A single strong system prompt no longer meets that standard on its own. For a writing pipeline specifically, every point where external content enters, brand corpora, SEO briefs, customer interview transcripts, needs to be treated as a potential injection vector and sanitized before it ever enters any model's context window.
The architectural approaches that address injection at the pipeline level rather than at individual agent boundaries
Defending a multi-agent writing pipeline against injection takes changes at the pipeline's architecture, not stronger system prompts bolted onto individual agents one at a time.
Boundary verification is the first and most fundamental piece: explicit validation of content, identity, execution intent, and state integrity at every single inter-agent handoff in the pipeline. The Bappy et al. research identifies the absence of exactly this primitive as the core structural gap behind all three unverified boundaries.
LLM Tagging, proposed in the Prompt Infection research, offers a second layer: combined with existing safeguards, it significantly cuts down the spread of an infection once it starts. Tagging gives each agent in the pipeline a way to tell its own instructions apart from content it received from an upstream agent, so it closes off the exact ambiguity that lets injected instructions pass as legitimate directives.
A third approach comes from research testing two competing multi-agent defense architectures directly against each other: a sequential chain-of-agents pipeline, where a dedicated guard agent vets every candidate output before it moves downstream, and a hierarchical coordinator-based system. The guard-based chain-of-agents approach described in the paper A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks hit 100% mitigation across every attack category tested, and it still preserved full functionality for legitimate queries. In that architecture, the domain LLM produces a candidate answer first, and a mandatory guard agent checks that answer for policy violations, attack indicators, and format compliance before anything is allowed to move forward. Only the guarded output ever reaches the next stage downstream.
Two more practical measures round out the defense. Audit-log streaming at every agent handoff, structured logs capturing session ID, tool name, tool input, and tool response, creates an evidence trail if injected content ever does manage to redirect a model's behavior, and it's practical to build into pipelines that already run on structured event contexts. And every ingested document, brand corpus material, SEO briefs, any external content a pipeline pulls in, needs to pass through a dedicated sanitization stage before it ever reaches a model's context window. The same skepticism a well-built pipeline already applies to user-supplied text has to apply to retrieved documents too, because the research is clear that the retrieved document, not the user's query, is where most of this risk actually lives.
Sources
- Image-Based Prompt Injection: Hijacking Multimodal LLMs Through Visually Embedded Adversarial Instructions
- Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
- [2608.00718] Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
- [2609.09604] Watermarks Without Verification: AI Text Watermarking After the EU AI Act
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks


