

Reading an AI design portfolio: evidence versus prompt output
Stop treating the screens as the evidence. There are 203 pipeline-generated product screens in the case studies on this site โ I built the pipeline and I direct it, and a convincing one costs me a command. Polish has come loose from the work that used to produce it, so a portfolio review that grades polish is now grading the wrong variable.
That number is from my own site, which is the only reason I'll say this as bluntly as I do. The generated studies here are labelled as concepts and I'm not passing them off as client delivery โ but they look exactly like client delivery, and that's the problem I'm describing rather than one I'm exempt from. I sell fractional design leadership, so I have an interest in you weighting judgment over artifacts. Read the argument, not my incentive.
What the number actually means
Fourteen studios feed the case studies here. The largest single study carries 48 rendered images; across the generated ones there are 203. They have consistent typography, a coherent token system, plausible empty states, real-looking data. A reviewer scrolling them cannot tell โ from the images alone โ which ones took a quarter with a client and which ones took an afternoon.
Neither can I, and I made them. That's the whole finding. The screenshot used to be a costly signal: producing forty polished screens meant someone spent weeks, which meant someone had made hundreds of small decisions, which is what you were actually trying to detect. The cost collapsed. The signal collapsed with it, and it did so faster than the review practices built on top of it.
What survived, and what didn't
| Still informative | No longer informative |
|---|---|
| What they discarded, and why | Visual polish, typography, spacing |
| The constraint they designed against | Screen count, project count |
| Where the work is wrong and they know it | Coherent design-system usage across screens |
| Who else was in the room, named | A tidy process narrative โ problem, research, solution |
| What shipped, and what happened next | Framework diagrams of any kind |
| The decision they'd now reverse | Case-study prose quality |
The right column isn't worthless โ it's cheap, which is a different claim. Cheap signals still separate the bottom of a pile from the middle. They no longer separate the middle from the top, and the top is presumably who you're hiring.
Notice that everything in the left column is a negative: the thing not chosen, the constraint that bit, the part that failed. That isn't a coincidence. Generative tools are overwhelmingly trained and prompted toward the plausible-positive โ a coherent artifact that looks like the finished output of a good process. Nothing in that machinery reliably produces "we tried the obvious thing for five weeks and it was wrong, and here's the residue."
The three questions to ask of the artifact
Before you spend anyone's time on an interview. Each one is answerable from the portfolio page itself, or conspicuously isn't.
- Where is the provenance? Not "did you use AI" โ a useless question in 2026, and one every candidate has a prepared answer for. The useful version: does the case study say who did which part, including the machine? "I specified the states; the layouts were generated against our token contract; I cut about a third" is a real answer. Uncredited solo authorship of an entire system is the tell, and it was the tell before AI existed.
- What was the binding constraint? Every real project is deformed by something โ a legacy schema, a compliance rule, a customer who pays for the module nobody likes, six weeks. A portfolio with no visible constraint is either a concept piece or a sanitized one. Concept pieces are fine, and the good ones say so.
- What did they throw away? Ask the portfolio, then ask the person, and compare. This is the single highest-yield question in the whole review and the hardest to fabricate, because a discarded direction has to be specific and worse in an interesting way.
Being fair to the existing red-flag lists
There is decent published work here and it deserves credit rather than being ignored so my version looks new. Verified Insider's portfolio red-flags piece and Learnist's list of AI case-study red flags both name real ones โ claiming a percentage improvement with no baseline, showing only what worked, taking solo credit for team output, presenting a Double Diamond diagram as though it were evidence of thought. The Fountain Institute's write-up on what gets senior designers hired lands the same point I'm making about depth beating volume.
Two honest limits on all of it. Most of it is written for the candidate โ how to make your portfolio look less inflated โ which means the heuristics are public and the sophisticated candidate has already applied them. And most of the specific red flags predate generative tooling: they detect inflation, a person overstating their contribution, which is a different failure from manufacture, where the artifact is real, coherent, and simply cheap.
The circulating screening statistics deserve the same scepticism I'd apply anywhere. You'll see claims that 78% of recruiters now use AI-assisted portfolio screening, that a review lasts two to three minutes, and that a homepage gets ten to fifteen seconds. I can't trace any of those to a published methodology or sample, and I'd treat them as folklore with a number attached rather than as measurement. The directional claim โ reviews are short and increasingly machine-mediated โ matches what I see. The decimals don't come from anywhere I can check.
What to do instead
Cut portfolio review down to what it's now good for: a twenty-minute screen that answers one question โ is there any evidence of judgment here โ and then stop. Do not run a second, deeper portfolio pass. The information you want isn't in the artifact any more, and a longer look at a cheap signal just produces a more confident wrong answer.
Then spend the time you saved on work you can actually watch: a real ticket in your real repo, tools allowed, with you watching how they turn the crank. That post is the method; this one is only the filter in front of it, and the whole sequence โ screen, interview, paid trial, references โ is here. If you're still deciding whether you want a person at all rather than a firm, that's a different question and it comes first.
When the portfolio is still enough
Two cases, and they're more common than the rest of this post implies.
High-volume IC hiring where consistency is the job. If you're staffing production design against an established system, the portfolio tells you what you need โ can they hold a system, do they have taste inside constraints. The judgment you'd dig for isn't what the role is buying.
Junior hiring. A junior portfolio was never strong evidence of shipped judgment; you're reading for potential and taste, and generated work is a weaker contaminant of that signal than it is of a senior one. Weight the discard question anyway, at whatever scale their experience allows.
Everywhere else โ anyone senior, anyone who'll set direction, anyone whose output you can't personally check every week โ the artifact is now the cheapest part of their application, and it should be the cheapest part of your review. The studies on this site are the proof, and I'd rather say so than let you use them as evidence of something they don't establish. How I'd rather be assessed, and what it costs; I'm at hello@product.inc. Navigate is the one I'd point at, and only because you can go read what happened to the team afterwards.