Brand consistency used to be a question of discipline: everyone reads the same manual, everyone follows it, and someone corrects the strays. Now that a growing share of first drafts comes out of language models, that mechanism no longer holds. A model does not open a PDF, it does not develop habits, and it keeps a rule only for as long as the rule sits in its context. If you want consistency, your style guide has to change: from a document written for people into a set of rules a machine can follow as well. Let us look at why the brand manual fails, where machine-readable brand rules are already standard, what models demonstrably forget, what a guide for both audiences looks like, and what the law adds in August.
Why does the brand manual fail as a PDF?
Almost every deck on the subject quotes the same figure: consistent branding lifts revenue by 23 per cent, or 33 per cent in the newer version. It is worth following that number back to its source. It comes from a Lucidpress press release dated 2 December 2019, which says verbatim: "consistent branding can increase revenue by 33%". The basis is a survey of more than 200 organisations, and the 23 per cent variant comes from the predecessor survey in 2016. That is self-reported data from marketing leaders, not audited revenue, and the older figure is now ten years old. Useful as a direction, not as the basis for a budget.
What the number obscures is more concrete and costs nothing to argue: the problem for most brands is not a lack of conviction, it is a rulebook in a format that never reaches the actual work. A 40-page manual is a reference work for people with time. It is not a rule a freelancer follows at 11pm, and it is certainly not a rule a language model follows. Consistency does not happen where rules are documented. It happens where they are present at the moment of writing.
Where does a machine-readable brand already work?
On the visual side the question is largely solved, and since recently there is even a standard. On 28 October 2025 the Design Tokens Community Group at the W3C published the first stable version of its specification, version 2025.10. Design tokens are nothing exotic: instead of "our blue is #324ceb" on page 12 of a PDF, the value lives in a file that Figma, iOS, Android and the website all read from. The organisations behind the format include Adobe, Google, Microsoft, Meta, Figma, Salesforce, Shopify and Pinterest. The announcement puts the purpose in one sentence: "The specification unlocks interoperability across design tools and code." In plain terms: one source of truth, many output channels.
What practice does with it is the interesting part, and it is sobering. The Design Systems Report 2026 by zeroheight ("This year, 147 design system practitioners shared their experiences", so a small specialist sample rather than a representative market) contains the sentence that matters: "Surprisingly, only 40% of teams have any kind of token pipelines established, which means they are manually syncing their tokens between design, docs and code." And only 54 per cent of teams had their tokens in the design tool, the code and the documentation at the same time.
That is the real lesson for the language side: a standard on its own does not produce consistency. Things become consistent only once the rule lands automatically where the work happens. For colours and spacing that pipe at least exists, even if six in ten teams have not laid it yet. For language most companies have nothing of the sort: the brand voice lives in a document meant for people and never seen by a machine.
Why does a language model forget your rules?
Everyone knows the drift from practice: the twelfth piece in a chat sounds different from the first, although nobody changed anything. Since May there are solid measurements for it. The SEQUOR benchmark, published on 7 May 2026, tests rule adherence not on single tasks but across 50-turn conversations, on ten open-weight models plus Gemini 3.1 Flash Lite.
The findings are uncomfortably concrete. Models hold a single rule across a long conversation reasonably well: accuracy falls by a good 11 per cent from the first turn to the last. As soon as three rules apply at once, accuracy drops by more than 40 per cent. And when rules are not stated upfront but introduced along the way, models lose 38 per cent on average. Even the strongest model in the test, Gemini 3.1 Flash Lite, lost 9 to 13 per cent depending on the regime. In fairness: the tested set is mostly mid-sized open-weight models, and the frontier models from the large providers should do better. No benchmark so far has contradicted the direction of the effect.
For everyday work that means three things. First, rules you slip in mid-conversation are the weakest form of instruction you have. Second, the more rules apply simultaneously, the less reliable each individual one becomes. A style guide with 40 commandments is not a stricter guide for a model, it is a blurrier one. Third, long sessions drift. Ask one chat to write eight pieces in a row and the tone at the end differs from the tone at the start, usually without anyone noticing.
How do you write a style guide a machine will follow?
The rebuild is smaller than it sounds. Four principles carry almost all of it.
Checkable rules instead of adjectives. "Friendly but professional" is not an instruction to a model, it is a mood board. "Second person, no exclamation marks, three sentences per paragraph maximum, every figure with a source" is one. The rule of thumb: if you could not test a rule with a short script, it is a direction rather than a rule. Both can live in the document, but keep them clearly separated so everyone knows what is negotiable.
Example pairs instead of definitions. Two sentences, one of them labelled "not like this", do more work than three paragraphs of brand philosophy. Models learn style from patterns, not from explanations. Three good pairs per rule are enough, and they double as the best onboarding material a new writer can get.
A hard banned list. The one part of a style guide that is 100 per cent machine-enforceable. Ours bans em dashes, filler phrases such as "game changer" and "revolutionises", generic triples, and any figure without a source. A banned list is unglamorous and still removes most of what makes text sound machine-made.
A rule budget instead of completeness. This follows directly from the benchmark above: keep the number of hard rules that apply at the same time small, with five to seven a realistic budget. Everything beyond that belongs not in the prompt but in the check afterwards. A model that reliably holds seven rules is worth more than one that vaguely knows forty.
And the most important point is organisational rather than editorial: the style guide does not belong in a chat, it belongs in the system prompt or in a versioned file that sits next to the content. That way it is reloaded on every run instead of explained once, every change is traceable, and all tools pull the same version. Precisely the pipe that the token pipeline provides on the design side.
How do you check that it holds, and what does the law require from August?
A rule nobody checks is a statement of intent. We run our own content production through a fixed gate for exactly that reason: before a piece is even uploaded as a draft, it passes deterministic checks against the banned list, the required structure and the sourcing rule, and it gets a score. Below the defined threshold nothing is published, it is reworked. That is unspectacular and effective precisely because it does not depend on how the day is going. What cannot be checked automatically, such as whether an example actually lands, stays with a person: human in the loop as a defined step rather than a slogan. Rules, ownership and sign-off taken together are simply content governance, the unglamorous part that decides whether a brand holds together as volume rises.
From 2 August 2026 a legal layer joins in, when the transparency obligations in Article 50 of the EU AI Act start to apply. Providers of generative systems must ensure outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated". For you as a deployer, the second part matters more: "Deployers of an AI system that generates or manipulates text which is published with the purpose of informing the public on matters of public interest shall disclose that the text has been artificially generated or manipulated." Then comes the exception that justifies all the work above: the duty falls away where the content "has undergone a process of human review or editorial control" and a natural or legal person holds editorial responsibility.
For context: the disclosure duty for text targets publications intended to inform the public on matters of public interest, not every product description. And under the Omnibus deal provisionally agreed in May 2026, generative systems already on the market before the cut-off get until 2 December 2026 for the machine-readable marking requirement. The practical consequence is the same either way, and it is good news rather than paperwork: the review step you need for quality anyway is the same step that creates editorial responsibility. Document your gate properly and you meet the exception as a by-product.
Three levers
1. One page instead of forty. Five to seven checkable rules, three example pairs per rule, one banned list. Anything beyond that is direction and belongs in a second section. That page is now the style guide, the manual stays background.
2. Lay the pipe. The same file in the content tool, the system prompt and the editorial process, versioned. The reason design tokens work is not the format, it is the pipeline. The same holds for language.
3. A gate before publishing. Check automatically what can be checked, and have a named human sign off on the rest. It keeps the brand together, and from August it is also your position on Article 50.
If you like, we can look at your existing style guide together: what a machine could actually follow today, what is only atmosphere, and which five rules to make hard first. 🙂
