Agree the measurement before the mock
Tags: experiments, metrics, instrumentation
Once a design exists, the measurement gets chosen to fit it â and the easy number wins. Page views. Total conversions. Time on page. Those move for reasons the change had nothing to do with, so the result is noise presented as evidence, and a feature whose success was never defined can be declared a success by whoever reads the dashboard most generously.
Why after is too late
The order matters more than it looks.
Before the variant exists, "what would count as this working?" is an open question with real answers. You can argue about it, disagree, and pick something the change could plausibly move. Nobody has anything invested yet.
After the variant exists, the same question has become a question about a thing someone made. The measurement that gets chosen is, almost inevitably, one that the thing does well on â not through bad faith, but because every metric has a story, and it is easy to find the story in which this design succeeded.
And the easy metrics are the dangerous ones. Aggregate page conversion is affected by the hero, the traffic source, the day of the week, a competitor's launch and a hundred other changes happening at once. A small true effect from one change drowns in it, and a random fluctuation looks like a win.
Scope the goal to what the change can move
The single most useful move is to measure the specific action the change touches, rather than the page's overall outcome.
If the change is a button label, measure clicks on that button â not every contact link on the page. If the change is how an AI suggestion is presented, measure whether that suggestion gets accepted â not session length. Narrowing the goal to what the variant can actually affect is what turns a comparison from a guess into a reading.
This is especially sharp for AI features, whose effects are often local and whose failures are often invisible in aggregate: a worse suggestion that people quietly ignore does not dent overall engagement, but it shows up immediately in the acceptance rate of that suggestion.
Make it structural
Good intentions about writing the goal first do not survive a deadline. The durable version is to make it impossible to skip.
If an experiment cannot be created without its goal, the goal cannot become an afterthought. Put it in the schema, the registry, the ticket template â wherever the experiment is born â as a required field.
That has a second benefit that is easy to miss. A goal declared at creation means every exposure is attributed correctly from the very first visitor. The numbers are being collected against the right definition from day one, whenever someone gets round to reading them â rather than a metric being chosen later and applied backwards to data that was never shaped for it.
Grounded in
This site's A/B registry makes goal a required field. An experiment cannot be registered without naming the event that counts as a conversion â the type will not let the entry exist otherwise.
The hero call-to-action test shows the scoping rule as well. It compares two labels on the main button, and its goal is not any contact link on the page. It is a click on that one button, tagged for exactly this purpose, so the number the two labels are compared on is the one a label can actually move. Measuring every mailto link would have folded in visitors who scrolled past the hero entirely and got in touch from somewhere else, which no change to that button could possibly have caused.
Anti-patterns
- Choosing the metric after the result. Every change succeeds by some measure, if you look for one afterwards.
- Aggregate metrics for local changes. Page conversion for a button label; a small true effect drowns in everything else.
- The goal as an optional field. Optional means skipped under deadline.
- Engagement as the goal for an AI feature. Time spent can rise because the output is confusing.
- Several goals, reported selectively. Measuring five things and presenting the one that moved is a fishing expedition with a slide.
The smallest version worth building
Before building a variant, write one sentence: this works if the rate of [specific action this change touches] goes up. Put it where the experiment is defined, and make that field required.
If the sentence cannot be written, the change is not ready to test â which is worth knowing before it is built rather than after.
Related patterns
- Evaluation loops â the offline counterpart: a fixed set of cases for judging a change before it reaches any traffic.
- Prototype the distribution, not the screen â deciding what to look for before reviewing applies to prototypes as much as to experiments.
- A cold product can't be judged â a realistic population is what lets a measurement be designed against real density before launch.