Runs need a shelf
Tags: history, comparison, versions
A reverse-chronological list answers what happened. The question people actually have is which of these is best β and that one cannot be answered from a list, because comparison requires adjacency and a list only ever gives you one thing at a time.
A list biases the outcome it was meant to record
A person can hold one artifact in working memory. Not two. So a list asks them to remember the one above while reading the one below, which nobody can actually do β and the resolution is always the same: they take the most recent.
That is the real cost, and it is worth stating plainly. You generated eight options and recency picked the winner. The model explored the distribution, the interface threw the exploration away, and the product reports this as a feature because all eight are still there if anyone wants them.
Nobody wants them, because wanting them would mean doing the comparison by hand.
What a shelf needs that a log does not
- Adjacency. Side by side, same screen. This is the whole pattern and everything else is support.
- A stable identity per run. "The third one" is not addressable. If a user cannot refer to an attempt, they cannot discuss it, bookmark it, or come back to it β and they certainly cannot tell you which one was good.
- The input visible alongside it. Without the prompt or parameters that produced it, the user is comparing outcomes with no access to the variable. See regeneration is an edit, not a retry β the resample and the edit need different things shown, and a shelf that hides the input collapses them again.
- An explicit act of keeping. Selecting a winner has to be a thing the user does, not something inferred from which one they stopped on.
The log falls out of that for free. Build the comparison and you have the history; build the history and you have nothing.
When a shelf is the wrong build
The bottom branch is easy to miss. A shelf of attempts is a shelf of data, and every attempt carries whatever was in its input. A product that accumulates forty generations per user per day is accumulating forty prompts per user per day, which becomes a retention question and sometimes a disclosure one.
Grounded in
User Study's session stack, built for moderators running back-to-back multi-angle research sessions β phone screen, workspace, and the assistant under test, recorded across a study.
The recordings are many attempts at the same shape of thing, and the design problem was never storage. It was keeping a study scannable across all of them: a moderator coming out of the fifth session needs to find the moment in the second where something went differently. That is a comparison requirement, it produces a different interface from a file list, and the session stack is what it produced.
Anti-patterns
- A reverse-chronological list presented as the comparison interface. The default, and it silently selects for recency.
- No stable identity per attempt. Unaddressable, so undiscussable.
- Attempts stored without their inputs. Comparing results with the variable hidden.
- Recency as the default selection. The interface making the choice and calling it the user's.
- No way to mark a keeper. So the best one is whichever was last, again.
- Unbounded retention by default. Every attempt is also every prompt.
The smallest version worth building
Show the last few attempts side by side, each with the input that produced it and a button that says keep this one.
That is the entire pattern. The log you already have; what is missing is the row.
Related patterns
- Regeneration is an edit, not a retry β where the attempts come from, and why a resample and an edit need the shelf to show different things.
- Generated things need generated names β a shelf of identically-named attempts is still unscannable.
- Evaluation loops β the same comparison discipline applied to the product rather than to one user's task.