Est.

Convergence Criteria in Iterative Writing Systems

Iterative systems need control loops that catch errors mid-process, not just stronger models.

Staff Writer · · 12 min read
Cover illustration for “Convergence Criteria in Iterative Writing Systems”
Adversarial Writing Systems · September 30, 2026 · 12 min read · 2,654 words

Single-pass LLM generation fails because the architecture forces researching, drafting, and evaluating into one pass with no mechanism to catch its own mistakes. A model asked to research a topic, write about it, and check its own work in one continuous generation has no point at which it can stop, notice an error three paragraphs back, and revise before continuing. There's no branch point, no retry, no internal critic. Whatever gets written in sentence one constrains everything that follows, and if sentence one contains a wrong assumption, the rest of the output inherits it.

The formal framing for this comes from work on agentic artifact creation surveyed by Wang and colleagues. Direct generation, in that framing, follows a fixed control schedule: the system commits to actions in sequence, and intermediate observations, the things it notices as it goes, have no way to redirect what comes later. Inspection only happens at the end, once the artifact is already finished, and by then a failure early in the sequence has already propagated through everything downstream. Once a fault is finally caught, it's hard to localize to its origin, so the fix is throwing out large sections and regenerating them rather than making a targeted edit. It's throwing out large sections and regenerating them.

This limitation appears even in systems with multiple agents involved, and the Emergence World study from Akkil and colleagues at Emergence AI documents goal drift as a related failure when agents never diverge from their assigned coordination: agents gradually diverging from the objectives they were assigned, with the drift accumulating across interactions rather than announcing itself in any single exchange. A one-off evaluation of a session in isolation won't catch that kind of slow divergence, because nothing in that session looks wrong on its own, and the problem only becomes visible across time, so the fix has to operate across time too.

The lesson from both threads is the same: single-pass systems need a control loop, not better prompting or a stronger model. It's a control loop, some structure that lets earlier output be checked and revised before the system commits to a final artifact. Without that loop, there's no way to define a stopping condition, because there's no mechanism to evaluate whether stopping is warranted in the first place. Everything the rest of this piece describes, the criteria, the review architecture, the voting, exists to fill that specific gap.

Stopping conditions and why prose is harder than code

A stopping condition is a definition, agreed in advance, of what counts as finished. In software engineering the acceptance criteria are binary and machine-checkable: a test suite either passes or it doesn't, and a compiler either produces valid output or throws an error. Prose offers no equivalent artifact. There's no compiler for tone, no test suite for whether a paragraph sounds like it came from a particular brand rather than from the internet at large.

Wang and colleagues' survey of agentic artifact creation makes the underlying difficulty explicit: complex deliverables are governed by several acceptance criteria at once, the validity of each is only partially observable, checks that do exist cover different criteria and often arrive at different times, and a check's result can go stale the moment the content underneath it gets revised. That is a precise description of what makes prose acceptance criteria harder to pin down than code's. A unit test doesn't go stale when you refactor a function's internals, as long as the function's contract holds. A tone check on a paragraph absolutely goes stale the moment an editor rewrites the sentence next to it, because tone in prose is relational: it depends on everything around it, not on a fixed contract.

That relational quality is what makes the criteria genuinely interdependent rather than a list to be checked off separately. Changing the content of a claim can shift the tone with it. Shift the tone and the perceived authority of the piece can shift too. Shifting the perceived authority can make a factual claim that read as confident on the first draft suddenly read as evasive on the second, even though the underlying fact hasn't changed. Satisfying one criterion can quietly violate another, so a stopping condition for prose can't be a single threshold crossed once. It has to be a state verified across several criteria simultaneously, and verified again after each revision, because a revision that fixes one criterion routinely disturbs another.

The four dimensions writing systems need to converge on

A production writing system needs at least four distinct convergence criteria, and failing any one of them is grounds for continued iteration regardless of how well the others are satisfied. They're the axes along which the loop keeps running until every one of them is satisfied at the same time.

Voice fidelity is a separate and more elusive target. Output needs to match a brand's specific habits, its sentence-length tendencies, its vocabulary, its stance on the topics it covers, its emotional register, rather than settle into the model's own default. That default has a name. A 2024 Stanford NLP benchmark called it "assistant-speak," describing language that is polished, hedged, and tonally neutral, which is exactly the register a brand voice is usually trying to avoid. Left unchecked, this produces narrative drift, a brand's persona diverging across models, and the research brief notes that this is an infrastructure failure, not a style complaint: high AI share of voice with low narrative fidelity erodes brand value.

Structural coherence is the third criterion: argument, section order, and internal cross-references all need to hold together across the full length of the piece, and this is the criterion most exposed to the context-window pressure described earlier, since long-form work is exactly where a single pass starts losing track of its own earlier commitments. The fourth criterion, style constraint satisfaction, is the most mechanical of the four and the closest thing prose has to code's linting step: banned phrases, readability targets, consistent terminology, accessibility rules. LLMs default to verbal tics, phrases like "The reality is…" as an opener or "In summary, by leveraging…" as a closer, and a shared, living list of banned phrases, updated regularly, is a practical mechanism for catching them. Some of this checking can now run automatically: flagging violations against a style guide, scoring a draft against a readability target, catching terminology that drifted between sections.

Why adversarial multi-agent review is the mechanism that makes convergence measurable

Defining four criteria doesn't accomplish anything if nothing in the system can measure output against them, since a production writing system needs all four satisfied at once, not just defined. A single model, generating and then grading its own draft, tends to agree with itself: whatever blind spot produced an error in the first place is highly likely to be present in the second pass too, since the first pass and the second are running the same underlying judgment. Adversarial review breaks that symmetry by putting a different vantage point on the evaluating side of the loop than the one that produced the draft.

Developer Stavros Korokithakis documented a working version of this architecture in a production coding pipeline in March 2026. His setup uses OpenCode as a harness, with Claude Opus acting as architect, Sonnet 4.6 as implementer, and a panel of reviewer agents drawn from Codex, Gemini, and Opus supplying critique. The specific choice to draw reviewers from separate model families, rather than running several instances of the same model, is the load-bearing part of that design: a critic built the same way as the writer inherits the writer's assumptions, and a critique built on shared assumptions isn't really an independent check.

The Emergence World research backs this up from a different angle. Agent populations built entirely from a single model showed conformity even when individual agents privately disagreed with the group's direction, one of five recurring problems the study surfaced under sustained operation. That's a strong signal that model diversity isn't a matter of preference in these systems, it's closer to a structural requirement. A pipeline that splits planning, drafting, and review across models suited to each task runs checks that carry independent evidence. It's the only configuration in which those checks carry independent evidence rather than an echo of the same underlying judgment repeated three times over.

How the loop knows when to stop: voting, thresholds, and the convergence signal

The stopping condition in a system built this way is the point at which disagreement among critics resolves into agreement across every one of the four criteria at once. In practice, each critic agent is assigned to evaluate the draft against one or more of those criteria and returns a verdict on its slice of the problem; the loop exits only once all four verdicts come back satisfied simultaneously, the moment any single critic signs off is not enough.

Voting across critics drawn from different models is what makes that exit signal trustworthy rather than arbitrary. When critics with genuinely different training histories and different weightings independently land on the same verdict, the odds that they're all missing the same thing drop sharply, in a way that one critic's approval alone can't offer. That's the direct payoff of the model diversity argument from the previous section: agreement only counts as a strong signal when the agreeing parties didn't start from the same assumptions.

Panel size is itself a design decision with real cost on both sides. Four to six agents is the practical sweet spot: beyond six, context window overhead and debugging complexity grow faster than any further gain in quality. Not every check needs a vote, either. Banned-phrase lists and readability thresholds are deterministic and can be evaluated with a simple pass or fail, so they function as hard gates sitting outside the voting mechanism entirely: a draft with a banned phrase in it fails, no matter how the critic panel votes on everything else.

Iteration count still needs a ceiling, but that ceiling functions as a safety valve rather than a convergence criterion. It's a safety valve. A system that hasn't converged after a defined number of passes has a problem that more iteration probably won't fix, and the correct response is surfacing the draft for a human to look at, not letting the loop exit automatically with output that never actually satisfied its own criteria.

What voice fidelity convergence requires that accuracy convergence does not

Factual accuracy has an obvious external reference: the sources a claim is checked against either support it or they don't. Voice fidelity has no comparable outside anchor. The only legitimate benchmark for whether something sounds like a given brand is that brand's own prior output. A critic checking for voice fidelity needs a corpus of that output on hand before it can render a verdict.

There are two established ways to supply that reference. Context injection provides a dense, structured brand brief at every single generation, spelling out the rules directly. Retrieval-augmented generation instead pulls from a curated corpus of the brand's past writing at generation time, surfacing examples that match the specific task at hand rather than restating abstract rules. RAG tends to suit a living voice better, since new examples can be folded in as positioning shifts without retraining anything underneath the system. Fine-tuning adjusts the model's weights directly using labeled training data, and requires GPU compute and engineering time at a cost well above the other approaches.

How large that reference corpus needs to be depends on what's being produced. Long-form work, blog posts, whitepapers, calls for a substantial, genuinely representative body of polished text to draw from. Short-form work asks for far less: five to fifteen strong examples usually cover it. Without that corpus, or with one too thin to be representative, the system produces a flattening effect: the output still converges on something, just not on the brand's actual voice, settling instead into a generic version of whatever category the brand happens to occupy. Sourcing that corpus isn't a purely technical decision, either.

Watermarking as a convergence-adjacent problem: what detection signals tell you

Watermark detection tells a system whether a given model produced a piece of text. It doesn't tell the system whether that text is any good, and treating a detected watermark as a proxy for quality is a category error that pushes a system toward optimizing for undetectability instead of for the criteria that actually matter.

The regulatory backdrop is no longer speculative. The EU's AI Act and California's AI Transparency Act both began requiring watermarking of AI output on August 2, 2026, with California's date pushed back from January 1, 2026 by AB 853, and California's mandate applying specifically to images, audio, and video rather than all AI output. Providers have moved at different speeds under that pressure. Google announced a SynthID-based detector in May 2025 and expanded its rollout in May 2026; Anthropic said in August 2026 that Claude would watermark its output, covering new models starting August 2, 2026 and extending to older models by December 2, 2026; OpenAI has said since 2024 that it built a watermarking system for ChatGPT but chose not to deploy it, citing concerns about false positives against non-native English speakers and the risk that flagged users would simply switch to a competing product.

A detected mark confirms that a model generated the text. It says nothing about whether that text is accurate, on-voice, or structurally coherent, none of which are things a watermark was ever designed to measure. The absence of a mark is less informative still: it might mean the text came from an unwatermarked model version, a different AI system entirely, or an actual human author, with no way to tell which from the absence alone.

The underlying mechanism, going back to the Kirchenbauer method, has the model's own generation process leave a statistical fingerprint behind as it writes. The right way to treat watermarking in a production pipeline is as a compliance constraint the system has to satisfy on its own terms, alongside style enforcement, never as evidence that the underlying convergence problem has been solved.

How security failures in multi-agent systems look like convergence failures from the outside

A writing system that has been compromised and a writing system that simply hasn't converged yet can produce output that looks identical from outside the pipeline. Both appear as a draft that fails one of the four criteria: a claim that doesn't check out, a tone that's drifted from the brand corpus, or a structural seam where two sections don't quite fit together. Nothing about the failed output itself tells an operator which cause is behind it.

That ambiguity matters because the review architecture described earlier, critics from different model families voting on a shared draft, is also an attack surface. If a single reviewer agent in a four-to-six agent panel returns compromised or manipulated verdicts, the effect on the voting mechanism looks exactly like the effect of an honest critic with a legitimate disagreement: the loop either keeps iterating past its normal point or exits on a draft that shouldn't have passed. A system reading only its own convergence signal has no built-in way to distinguish a critic that's wrong from a critic that's been tampered with, because both produce the same symptom: a vote that doesn't match the quality of the underlying draft.

This is precisely why the safety-valve principle from the stopping-condition discussion carries weight beyond ordinary quality control. An unresolved factual dispute among agents and a compromised reviewer skewing the vote both produce the exact same signature: a system that keeps failing to converge after its defined iteration ceiling, or that converges suspiciously fast on output that a human reviewer flags as off. Treating every convergence anomaly as purely a quality problem, rather than as a signal that might warrant a security check on the agents casting votes, leaves that second failure mode invisible until it's already shipped.

Sources

  1. Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
  2. Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
  3. AI watermarking - Wikipedia

More in Adversarial Writing Systems