Building a Brand Corpus for AI Voice Training
Ditch adjectives and build AI voice training from actual published writing patterns instead.

Most brand voice work fails before a single word of content gets drafted, because the document it hands to an AI system is a list of adjectives, and adjectives describe intent rather than behavior. To a model, it functions as a prompt to produce the statistical median of everything in its training data ever tagged with those two words. That median is nobody's voice. It's the blurred average of thousands of brands that each reached for the same safe pair of descriptors. That is why so much AI-drafted marketing copy sounds like it was written by the same anonymous hand regardless of which company's name sits at the top.
The deeper issue is that terms like "bold but approachable" describe nearly every brand that has ever briefed a copywriter, and so they constrain none of them. GOV.UK's own content-design guidance makes a version of this point directly: it keeps official prose deliberately emotionless and warns that adjectives can be subjective and make text sound more like spin than information. The adjective is the heading. The rule is the actual instruction.
Mailchimp's content style guide, widely used as a reference across the marketing industry, follows the same architecture. It separates voice, which stays constant, from tone, which shifts by context, and it operationalizes voice into four named values with explicit rules sitting under each one. Named values sit at the top, mechanical rules sit underneath, and worked before-and-after examples sit at the bottom of this structure. None of them stop at adjectives, because adjectives alone were never built to survive contact with a drafting process, human or machine.
The usual objection is that a marketing team already has brand guidelines, complete with examples of approved copy. Those examples don't function the way teams assume they do. A model can see that an approved sentence runs twelve words and still have no way to know whether twelve words is the pattern or an accident, because nothing in the document told it which parts of the example are the rule and which parts are incidental. That gap between showing and specifying is the problem the rest of this piece solves.
How revealed behavior differs from stated intention
A brand guidelines document states intention. A published corpus of a brand's actual writing reveals behavior, and only one of those two things is something a machine can act on reliably. When a document says "we are direct and human," a model has no way to test whether a given sentence satisfies that claim. It can only treat the words "direct" and "human" as statistical pointers and search for output that correlates with other writing described the same way, which tends to produce something generically casual rather than specifically direct in whatever way this particular brand is actually direct.
A corpus of real, published work doesn't make claims about itself. It contains patterns that can be measured: the shape of sentence-length distribution across a hundred posts, the ratio of hedging language to assertive language, the specific words the brand reaches for again under deadline and the words it never touches. These are falsifiable. A reviewer can check whether a new draft matches the pattern, and a linter can enforce the match automatically, because the pattern was extracted from evidence.
Exemplars help more than adjectives do, but they don't close the gap by themselves. A peer-reviewed study evaluating six frontier language models tested authorship verification on informal blog writing and found accuracy falling as low as 17 to 21 percent, even when the models had real writing samples from the target author available to them. If models with direct examples of an author's actual sentences still can't reliably tell that author's writing apart from an imitation, then handing a model a folder of "good examples" and trusting it to infer the underlying pattern amounts to a hope dressed up as a strategy.
The distinction between stated intention and revealed behavior carries a cost that compounds over time. Voice has to be extracted from what a brand has actually published, because what a brand believes about its own writing and what that writing actually does are frequently two different documents.
Corpus hygiene as a design decision, not a cleanup step
What gets fed into the extraction process shapes whether the system learns the brand's standard or its average, and that choice belongs to an editor, not a script. None of that is wrong exactly, but none of it is the standard either, and a system trained on the average will produce average output.
Feed a curated set of the brand's edited best work instead; the extraction produces the standard, the patterns present in the writing the team would actually hold up as representative if asked. The resulting set should stay compact and current. A smaller pile of material everyone agrees is representative extracts cleaner signal than a sprawling archive full of internal disagreement about which parts count.
Corpus hygiene isn't a one-time pass either. The most common failure here is treating the live website as the corpus by default, on the assumption that whatever is published must be correct. None of that disqualifies the website from contributing examples. It does disqualify the website from being trusted wholesale without the same editorial filter applied to everything else.
Stylometric measures worth computing
Only a handful of stylometric measures are stable enough to serve as an actual specification, and knowing which ones to skip matters as much as knowing which ones to run. Paragraph rhythm is worth tracking too, since how a piece tends to open an argument, how it builds across a paragraph, and where it lands at the end are measurable traits.
MTLD, the Measure of Textual Lexical Diversity, is normalized for length. It stays stable whether you're comparing a short post against one several times longer. Hedge-to-booster ratio, counted per 100 words using a word list built for marketing prose specifically, is worth computing because it captures something real about a brand's posture: how often it qualifies a claim with words like "may" or "often" versus how often it asserts flatly. That ratio is one of the clearest fingerprints a brand has, and it rarely appears in a guidelines document.
None of this computation needs specialist tooling. Any scripting language handles sentence-length distributions, MTLD, and hedge ratios without difficulty. These measurements go missing from most brand voice work because almost nobody runs them. Vocabulary boundaries belong in the same spec and are just as measurable: which words the brand uses constantly, which it avoids entirely, which it reserves for narrow contexts. Those boundaries can be written as plain include and exclude lists. The exclude list carries as much weight as the include list: ban corporate jargon like "synergy," "leverage," and "transformative," and ban the stock phrases generic models reach for under pressure, things like "in today's fast-paced world," "it's important to note," and "in conclusion." These patterns are predictable enough that banning them outright costs nothing and prevents a specific, common failure mode.
Structuring measurements into a spec the system can follow
Raw measurements don't produce consistent output on their own. They have to be formalized into a structure the system can actually follow, and the structure that works has three layers: named values, mechanical rules under each value, and worked examples under that.
Layer one is voice identity: what the corpus reveals about sentence architecture, vocabulary boundaries, rhythm markers, and stance, including a default confidence posture the brand tends to write from, whether that's certain, exploring, or questioning.
Counterexamples earn a place in this structure too, and they carry as much weight as the positive examples. An actual off-brand passage, with a short note explaining precisely why it fails the spec, is something a system can test against directly.
The deliverable that comes out of this process should stay small: a worksheet of around seven rows, each one pairing a single measurement with the rule it produced, compact enough to fit inside a system prompt and survive contact with an editorial team that has no patience for a forty-page style guide. One more structural discipline matters here: keep facts, voice, and instructions in three separate buckets. It should leave the voice guidance untouched. This spec is a machine-actionable translation of what the corpus actually shows, built to be reviewed rule by rule rather than judged only by whether the final output happens to sound right.
Why a single model drifts when it enforces its own voice rules
A spec this precise still fails if it's handed to a single model and that model is asked to both draft against it and check its own work. The same statistical tendencies that pull a model toward an off-brand construction in the first place are active during its self-review, so the model has no real way to catch what it just produced. It drifts, and it approves its own drift, because both steps are running on the same underlying weights.
The fix is architectural, requiring a change to the system design itself. The generator focuses only on producing a draft. Research on multi-agent critique pipelines backs this structural choice directly: these pipelines exist specifically to mitigate the flawed critiques a single model tends to generate, aggregating feedback from multiple agents rather than trusting one model's judgment of its own output.
The fix is role specialization, not better prompting of a single model. One model drafts. A second, distinct model or layer enforces tone and voice. That second model can only enforce what the spec gave it to test, which is the exact reason the earlier work of turning adjectives into mechanical, checkable rules matters. It can enforce "sentences average under 18 words" and "no instance of the word synergy," because those are rules with a true-or-false answer. This is the architectural conclusion the whole argument has been building toward: a voice spec only holds up in production once it's checked by something other than the model that wrote the draft.
Encoding the spec in the harness, not just the prompt
A spec that lives only inside a system prompt stays fragile, because prompts get edited, forgotten, or quietly overridden by whoever touches the pipeline next. A spec encoded into the surrounding infrastructure, the harness that wraps the model, holds up far better across agents and across sessions.
The harness concept comes from software engineering, and it's the right frame for this problem. The model itself handles reasoning and generation. For writing specifically, repository-level configuration files, an AGENTS.md and optionally a CLAUDE.md, enforce style consistency across every agent that touches the repo, defining style rules, naming conventions, directory structures, and frontmatter requirements in one place that every agent reads before doing anything else.
This infrastructure now operates at meaningful scale. Agents don't tolerate the kind of inconsistency a human developer might shrug off. The rules governing their behavior need to live somewhere more durable than a prompt.
Passive documentation gets ignored. What changes behavior reliably is a hard steering signal placed directly in the loaded context: a skill file, an error message, a CLI prompt the agent has to respond to before continuing. Voice rules encoded as linter checks work the same way code quality checks already do in a software pipeline: a failed check produces a machine-readable error, and that error gets fed back into the agent's context as a remediation instruction, not as a suggestion buried in a document nobody reopens. This is the point where developer tooling and content production converge concretely. The voice spec belongs in the repository next to the code, checked with the same tooling discipline applied to any other output standard the team enforces.
The only honest acceptance test for a voice corpus
None of the extraction, structuring, or harness engineering described above proves anything on its own. The only acceptance test that actually validates a voice corpus is blind discrimination: whether people who know the brand well can tell generated content apart from real published content when neither is labeled.
The protocol is direct. Ask them to sort real from generated, then compute how often they get it right.
The scoring here is specific. Under the Unpromptable Test framework, attribution accuracy below 70 percent means the voice training is functionally useless. Between 70 and 80 percent, the spec is close but still leaking somewhere identifiable. Above 80 percent, the prompt architecture and the underlying spec are working. That threshold turns "does this sound right" from a matter of opinion into a number a team can act on, and it turns failure into something diagnosable rather than just a vague sense that the copy feels slightly off. When reviewers misattribute a piece, the reason they give for their guess points directly at which measurement in the spec needs tightening, which makes this test an ongoing process. The corpus and the spec built from it stay accurate only as long as that test keeps running.


