Prototype the distribution, not the screen
Tags: prototyping, review, variance
A mock freezes one sample, and it is always a good one, because somebody chose it. Reviewers approve the best case. Then the feature ships into the whole distribution â and none of what goes wrong there was visible in the artifact the decision was made on.
A mock of generated output is a picture of the best case
In deterministic software, a mock is a reasonable proxy for the shipped screen. The same input produces the same output, so the thing you reviewed is the thing people see.
A generated-output product breaks that. The answer in the mock is one sample out of a distribution, and it was selected â usually unconsciously â for being a good example: the right length, well-formed, on topic, pleasant to look at. The distribution the feature actually meets includes:
- the answer three times longer than the card was designed for,
- the empty answer,
- the refusal,
- the confidently wrong answer that looks exactly like a right one,
- the one in a language nobody tested.
None of those can appear in a static mock, so none of them get reviewed. The review approves a screen that represents the product's best moment and says nothing about its typical one.
Make the deliverable something that runs
The reviewable artifact for a generated-output feature is a working prototype â something that takes real input and produces real output through the actual interface.
And then, crucially, run it across a spread of real inputs, not a hand-picked few. Twenty varied cases reviewed together say more about a feature than one perfect case polished for a meeting. The questions change from does this look good? to what does this interface do when the answer is long, or empty, or wrong? â which are the questions that decide whether it works.
Turn assumptions into checks
A running prototype has a second advantage a mock can't offer: it can assert things.
A mock implies that a state appears â after submitting, the confirmation shows. A prototype can check it, and fail loudly when it doesn't. Once the prototype carries its own assertions, review stops depending on who happened to click through what, and becomes something repeatable: the same flow, the same checks, run again after every change.
That is where prototyping and evaluation meet. The prototype runs the distribution; the assertions mark what must hold across it.
When a static mock is still right
- Before there is any output to run. Early exploration of layout and flow does not need a live model.
- For the parts that don't vary. Navigation, settings and chrome are deterministic; mock them freely.
- For communicating intent to people who won't run it. A mock is a fine illustration, as long as nobody mistakes it for evidence.
Grounded in
The flow recordings on this site come from a running prototype rather than an animation. A real cursor drives the actual interface through each flow, and the recording carries expect assertions â this screen should be showing, this text should be on screen, this should now be gone â so a failed expectation ends the recording at the step that broke. Every recording is then rendered deterministically to video.
That makes each recording evidence that the interaction works, not an illustration of how it is supposed to. A flow that silently stopped matching its design cannot produce a clean recording, which is precisely the property a hand-built animation does not have.
Anti-patterns
- Reviewing generated output from a single mocked example. Approves the best case and nothing else.
- Cherry-picked demo inputs. A prototype run only on the cases known to look good is a mock with extra steps.
- No long, empty or wrong cases in the review. The failures that define the feature, left out of the artifact that decides it.
- Prototypes without assertions. Every review becomes a fresh, unrepeatable click-through.
- Treating the illustration as evidence. A beautiful animation of the intended flow proves nothing about the real one.
The smallest version worth building
Wire the prototype to the real model. Collect twenty real inputs, deliberately including a long one, an empty one and one that should be refused. Review the interface's behaviour across all twenty in one sitting.
Then add assertions for the three states that must always appear. That is the difference between a demo and a check.
Related patterns
- Evaluation loops â the fixed set of real cases this prototype should run, kept stable over time.
- Agree the measurement before the mock â decide what the review is looking for before seeing the output.
- A cold product can't be judged â a realistic population is the input spread a data-shaped prototype needs.
The full argument: How I work: the deliverable is something running.