Decide what counts as the same request
Tags: caching, cost, mental models
Every generative product answers the question is this the same ask as last time? â usually three times, in three different places, without anyone noticing they are the same question. When the three answers disagree, the product produces surprise bills and surprise outputs, and both come from one unmade decision.
Three definitions of sameness
The cache has one. Somewhere there is a key. It might hash the full prompt, or the user's inputs, or a document ID plus a version. Whatever it includes decides what gets reused and what gets regenerated.
The cost model has one. Whatever is not reused is paid for â in money, in latency, in rate limits. So the cache key silently determines the bill.
The user has one. Every person carries an expectation about whether a small edit is a small change. Fixing a typo in a prompt feels like nothing. Reopening yesterday's work feels like it should be instant.
These are usually defined by different people for different reasons. The cache key is whatever was convenient to hash. The cost model is a consequence of the cache. And the user's expectation is never written down at all.
Where they disagree, things go wrong in both directions
- The key is stricter than the user's expectation. A one-word edit changes the hash, so the product regenerates the whole job â slowly, expensively, and with an output that now differs in ways the user did not ask for. They thought they made a small change; they got a new artifact and a bill.
- The key is looser than the user's expectation. Two requests the user considers different hash the same, so they are served a stale result for something they believe is new. This is the more dangerous failure, because nothing tells them it happened.
Both of those are the same bug: sameness was defined by accident.
Define it once, in the user's terms
The fix is to decide deliberately, and to decide it in terms the user would recognise.
Start from the question what would this person call the same request? â then derive the cache key from that answer, and let the cost model follow from the cache. That order matters. A cache key chosen first, for engineering convenience, forces the user's mental model to adapt to your hashing scheme; a definition chosen first, in their terms, makes the cache behave the way they already expect.
Then make it legible. Before an action that will regenerate, the user should be able to tell that it will â and, where it costs something noticeable, roughly what. The worst outcome is not a regeneration; it is a regeneration the user did not know they were triggering.
Grounded in
Pathfinder keys its render cache on the full loadout, so the app pays to generate each distinct look exactly once.
That single choice makes the loadout â not the button press, not the session, not the order in which items were equipped â the unit of sameness. Equip a combination you have worn before and you get that exact portrait back, free and instant, because by the product's own definition you asked for the same thing. The cache key, the cost and the user's intuition all agree, because the definition came from the thing a player recognises as their character's look.
Anti-patterns
- Hashing whatever was convenient. The cache key ends up defining sameness, and nobody chose it to.
- Invisible regeneration. An edit that silently re-runs the whole job, with the cost discovered later.
- A key looser than the user's intent. Stale results served for requests the user believes are new â the failure nothing surfaces.
- Order-sensitive keys. The same selection made in a different sequence counted as a different request.
- No way to force a fresh result. When a user genuinely wants a new attempt, the cache should not be able to overrule them.
The smallest version worth building
Write down, in one sentence, what a user of your product would consider the same request. Compare it to your actual cache key. Every difference between the two is either a surprise bill or a stale result waiting to happen.
Then add one line of interface near any action that will regenerate: this will create a new version. It costs nothing and removes the most common version of the surprise.
Related patterns
- Hold the invariant, vary the delta â what is held fixed across a change is part of what makes two requests the same.
- Regeneration is an edit, not a retry â the other side of this decision: what happens when the user deliberately asks for something new.
- Runs need a shelf â once repeated requests produce distinct results, those results need somewhere to be compared.