Generation touches tokens, never screens
Tags: design systems, tokens, consistency
Ask a model for screens and you get work that is individually plausible and collectively incoherent: each screen fine on its own, the set drifting in spacing, colour and hierarchy, with no single place the drift can be checked or corrected.
The screen is the wrong unit to generate
A generated screen is a self-contained decision. Its spacing, its type scale, its colour choices were all made for that screen, in that moment. The next generated screen makes those decisions again, slightly differently.
No individual screen is wrong. The problem is the set. By the tenth screen, there are ten spacing scales that are nearly the same, several shades of the brand colour, and headings that are almost but not quite consistent â and none of it can be fixed in one place, because it was never decided in one place. Consistency has to be audited after the fact, screen by screen, which is the expensive version of design-system work that design systems exist to prevent.
Point generation at the system instead
The alternative is to let a model change the variables a deterministic layout reads â tokens, component parameters, the scale itself â and leave composing the screen to code that cannot improvise.
The difference is leverage and verifiability:
- A token change reaches every screen by construction. It is decided once, reviewed once, and applied everywhere it is referenced.
- It is checkable in one place. Whether a proposed palette meets contrast, whether a spacing scale is coherent â all of it can be verified before anything renders.
- Screens stay consistent automatically. Composition is deterministic, so two screens using the same components with the same tokens cannot drift from each other.
So the model does the part it is good at â proposing a coherent system â and never touches the part where its variability is a defect.
The generative UI objection
There is a real, opposite position: that models should generate interface directly, per request, because a UI shaped to each moment beats any fixed system.
It is worth taking seriously, because there are cases where it holds â a one-off visualisation of a specific answer, a throwaway view that exists for one question and is gone. When a screen genuinely has no siblings, there is no set to drift.
But most product surfaces are not like that. They are lists, forms, detail views and dashboards that users return to and expect to behave the same way twice. For those, per-request generation trades the consistency users rely on for a novelty they did not ask for â and it produces exactly the incoherence described above, at the speed of every request.
Grounded in
This site's design system. When counted for the long-form argument, 139 token names carried 4,138 references across the codebase.
That ratio is the whole case. A change at the token layer has roughly thirty-to-one leverage and is verifiable in a single file. A generated screen has one-to-one leverage and is verifiable nowhere. The design pipeline that produces the screens in the case studies here is built on that split: generation proposes and adjusts the system, and deterministic composition renders every screen from it â which is why the screens agree with each other without anyone checking that they do.
Anti-patterns
- Generating styled screens one at a time. Individually fine, collectively incoherent.
- Raw values in generated output. A model emitting hex codes and pixel values bypasses the system entirely.
- A token layer the model can't see. If the system is not legible to the generation step, it will reinvent it.
- Auditing consistency after generation. Checking ten screens for drift is the cost the token layer was supposed to remove.
- Treating every surface as a one-off. Generative UI earns its place on genuinely singular views, not on the list page users open every day.
The smallest version worth building
Give the generation step the token contract and forbid it from emitting anything the contract doesn't name. Anything it wants that isn't a token becomes a proposal to the system, reviewed once, rather than a value baked into one screen.
Related patterns
- Spend the novelty budget on the model â the product-level version of the same restraint: stock components everywhere except the new idea.
- Prototype the distribution, not the screen â a system that composes deterministically is also one you can run across many inputs and review.
- Evaluation loops â changes at the token layer are exactly the kind worth checking against a fixed set of screens.
The full argument: Why generation touches tokens, never screens.