๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.๐Ÿšง Designers never finish their own portfolio. This one ships rough on purpose and gets better in public. If something looks half-done, it probably is, and I'm on it.
Open menu
Switch to Darkhello@product.inc
Recording a product demo with Claude โ€” cover sketch

Recording a product demo with Claude

ยท Alex Zapadenko

A demo video is the most trusted artefact a product team ships and the only one nothing checks. Write the demo as a script instead of performing it, let every step declare what the screen should do, and a demo that has stopped being true fails instead of playing. There are five of these on this site, 64 steps, 24 of which carry a claim.

I sell design leadership, and the recorder here is mine, so read it as interested. The mechanism is small enough to describe completely, and the skill file is free.

Why demos rot without anyone noticing

A screen recording is one person's session with a wall clock running. It captures what happened once, on that machine, at that frame rate, including the moment they hesitated over the menu.

Nothing in it is checked. That's fine on the day it's made, and it's the whole problem three months later: the product changed, the recording didn't, and it keeps playing exactly as smoothly as before. There's no failure โ€” a stale demo looks identical to a current one. So the person who notices is a customer watching you claim something the product no longer does.

Teams handle this by re-recording before anything important, which means the demo is only as fresh as the last time someone remembered to be anxious about it.

A tape is a script of what a person does

The alternative is to stop performing the demo and write it down. Here a recording is a tape: an ordered list of the things someone does, and nothing else.

{ click: { text: 'Buy Now' }, why: 'the only call to action on the card' }
{ type:  { placeholder: 'Search' }, text: 'amazonite' }
{ press: 'Enter' }
{ settle: 400 }

Two details in that carry most of the weight.

Targets are named the way a reader sees them โ€” { text }, { label }, { role, name }, { placeholder } โ€” never a class name or a test id. A tape that says click the thing labelled Buy Now breaks when the button's label changes, which is exactly when a demo showing the old label should break. A tape addressed by .btn-primary-2 keeps working while showing something nobody can find.

The verbs are what a person does, not what a machine does: click, type, press, scroll, hold, settle. There's no "set state" and no "navigate to" โ€” if the only way to reach a screen is to fake it, the demo is claiming a path the product doesn't have.

The rule that makes it evidence

Any step may declare what the screen should do afterwards.

{ click: { text: 'Pay now' },      expect: { screen: 'confirmation' } }
{ click: { text: 'Amazonite' },    expect: { text: '17 specimens' } }
{ click: { text: 'Clear' },        expect: { gone: 'Amazonite' } }

A failed expectation ends the recording and names the step that broke. That is the whole difference between a video and a piece of evidence. Every "clicking this does that" in these tapes was true the last time the file was rendered, because the render fails otherwise.

Across the five tapes on this site there are 64 steps and 24 of them carry a claim โ€” roughly one assertion every two and a half actions. Not every step needs one; a scroll usually doesn't. The ones that matter are the transitions, because those are what a demo is actually asserting.

TapesStepsSteps with a claim
EarthWonders โ€” storefront, filters, checkout33910
DocRobot โ€” redact first, share at the exit22514

DocRobot's ratio is higher for a reason worth noticing: its tapes are about a confidentiality boundary โ€” what leaves the machine and when โ€” so almost every step is making a claim somebody could be harmed by if it were wrong. The assertion density follows the stakes, not the length.

It is not a screen capture

The recorder drives a virtual clock and steps the app one frame at a time: advance the world to frame N's time, pull every running transition to that moment, screenshot, repeat. Same tape in, same bytes out, regardless of how loaded the machine was.

That matters for one practical reason. A deterministic recording can be diffed. Re-render after a change and any difference in the output is a difference in the product, not in how the laptop felt that afternoon. A wall-clock capture can't tell you that, because two takes of the same session never match.

It's also why the pointer is real. Input goes through the browser's actual mouse and keyboard at viewport coordinates, so the page receives trusted events โ€” not el.click(), which skips every hover state and any handler that checks whether a human did it. The cursor you see drawn in the video is at those same coordinates on the same clock, so what the video shows and what the app received cannot drift apart.

This is a different job from animating a design, where keyframes are the source and the artefact is a loop. Here the running product is the source and the artefact is a test that happens to be watchable.

What Claude does, and what it doesn't

Same split as the audit and the design-system check.

Claude writes the tape. Given the running product, it can find the controls, work out the order, and write the steps โ€” including the assertions, which is the part people skip when they write these by hand because the demo already works today.

Claude cannot decide what's worth showing. Which flow is the demo is a judgement about what the audience doubts. A model asked to record the product will faithfully record all of it, in the order the navigation happens to be in, and produce four minutes nobody watches. The useful version is ninety seconds answering the one question a buyer actually has, and choosing that question is not a scripting problem.

Claude cannot tell you a passing tape is a good demo. The assertions prove the product does what the video claims. They say nothing about whether the claim is interesting.

Where this stops

Three honest limits.

Related: motion as data for the other render target, how I run a UX audit with Claude, the deliverable is something running, the free skill file, and the rate card.

Counts are from this repository's tape files on the date above, and will change.