

Checking a design system with Claude
Three numbers tell you whether a design system is real: how many of its names are actually used, how many places in the code escape it, and how many of those escapes render wrong. All three are countable, and a model will count them for you in a few minutes. On this site today it's 138 of 143 names, 249 escapes, and one that renders wrong.
The catch is the part nobody says out loud: those counts are only exact if the vocabulary is closed. That condition does most of the work, and it's what separates a measurement from a percentage somebody made up.
I sell design-system audits, so treat the argument as interested. The counting half is a free skill file โ the same one I run.
"We're on the design system" is a claim, not a number
Ask a team whether the product uses their design system and you'll get a confident yes. Ask what fraction of the interface it covers and the room goes quiet, because almost nobody computes it. The system exists, it's documented, there's a Figma library and a package โ and none of that is evidence that the code reaches for it.
The gap matters because the failure is invisible from the inside. Every screen someone built by hand looks fine on its own. What it costs you shows up later, all at once, when a brand changes or a mode flips and 4% of the interface doesn't follow.
So the check is not "do we have a design system." It's "how much of this product is actually wired to it, and where isn't it."
The three numbers
Coverage โ names used over names defined. The contract here holds 143 names; 138 of them appear somewhere in the code. That ratio tells you whether the vocabulary is the right size. A system where half the names are unused has been designed for a product nobody built.
Reach โ total references. 8,244 on this site. This is the leverage number, and it's the one that justifies the system's existence: it's how many places a single token edit lands. Divide it by the names and you get about sixty uses per name, which is the multiplier you're buying.
Escapes โ places the code stepped outside. 249 today. This is the number that gets quoted as "249 issues," and doing that is how audits earn their reputation. 248 of them are arbitrary values flagged for a human to glance at. One of them renders wrong: a raw shadow colour in a lightbox component that sits outside the contract and won't follow a brand when the theme changes.
One defect, 248 things to look at and probably leave alone. Both facts come out of the same command, and reporting only the first number would be dishonest in one direction while reporting only the second is dishonest in the other.
The condition that makes it exact
Here's the part that decides whether any of this is measurement or decoration.
The check on this site is exact matching, not pattern-guessing, and it can be because the vocabulary is closed. Every token in the contract expands to a known, finite set of utility classes. So a class in the source either is one of those strings or it isn't, and "off-system" is a decidable question rather than a judgement call. The default Tailwind palette is switched off here, which means an off-contract class emits no CSS at all and fails loudly instead of subtly.
Now take the same check to a codebase where the defaults are on, arbitrary values are ordinary, and a raw hex in a styled-component is nobody's alarm. Nothing is decidable any more. A tool can still produce a number, and it will look just as precise โ but it's now an estimate dressed as a count, and you have no way to tell which of the two you were handed.
So the first question to ask of any design-system coverage number is what made it decidable. If the answer is "the linter has a list of allowed values and everything else fails," the number is real. If the answer is a regex over hex codes, it's an approximation, which is fine as long as nobody rounds it to a percentage and puts it on a slide.
This is also the practical order of work. If you want the measurement, close the vocabulary first โ turn the defaults off so unsanctioned values break visibly. That's a day of work and it converts every future audit from an opinion into arithmetic.
The numbers move for reasons that aren't about the system
I ran the same check three times this month, and the third run is the instructive one.
| Sep 7 | Sep 14 | Sep 28 | |
|---|---|---|---|
| Files scanned | 465 | 492 | 532 |
| Components | 466 | 493 | 533 |
| Screens | 111 | 127 | 98 |
| Token references | 8,439 | 8,915 | 8,244 |
| Names in use | 138 of 143 | 138 of 143 | 138 of 143 |
| Escapes | 247 | 255 | 249 |
Files and components went up all month. References went down between the 14th and today, and screens fell by twenty-nine. Read as a trend line, that says design-system adoption dropped โ and it didn't. I deleted three case studies. The screens went with them, and their token references went too.
That's the trap in treating any of these as a metric. The denominator is the codebase, and the codebase changes for reasons that have nothing to do with the system. What survived all three runs unchanged is the ratio โ 138 of 143 โ and the one defect, which is still there, because once it was prioritised it sat below things worth more.
If you're going to track one of these over time, track coverage and the defect count. The raw totals are inventory, not progress.
What the count can't tell you
Five names are defined and never used. A tool reports that as five dead tokens and, asked to tidy up, will happily propose deleting them.
Four of the five belong to a namespace that brands inherit rather than declare โ it's optional by design, and removing it would be the audit doing damage. The fifth is a real orphan. Nothing in the code distinguishes them, because the distinction lives in why the namespace exists, and that was never written down anywhere a scanner can read.
That's the general shape of it. The measurement is genuinely free now, and the judgement about what the measurement means is the thing that isn't. Same split as the UX audit, for the same reason.
BlackLine is what the other side looks like when it works: a sprawling enterprise product being migrated to React without stopping, where the design system โ Pantheon โ was established as the single source of truth so design and engineering built against the same tokens. Nobody needed a coverage report to know it was working. They needed one to know where it wasn't yet.
When the number is lying to you
Three cases worth knowing before you quote your own figure.
- A high ratio on a small contract. 30 of 30 names in use is not better than 138 of 143; it may just mean the vocabulary is too small and the gaps are being filled with one-offs that the check can't see because nothing declared them off-limits.
- Escapes counted without severity. A count that mixes "renders wrong" with "arbitrary padding value someone should glance at" is one number doing two jobs. Separate them before anyone reacts to it.
- A percentage with no denominator stated. "94% design-system coverage" is meaningless until you know what was in the scan. This check deliberately excludes generated content and the studio surfaces, where off-contract colour is sanctioned. Excluding them is right; not saying so wouldn't be.
Related: what a design system audit costs, how I run a UX audit with Claude, why generation touches tokens, the free skill file, and the rate card.
Every count above came from this repository's own usage-graph check on the dates stated, and will change.