.comThis is falkster.com, the notebook. Falkster.AI is the company.Go to falkster.ai
ExecutionNew·Falk Gottlob··7 min read

Polish Stopped Being Evidence

Fidelity used to be a proxy for conviction in design review. AI broke that proxy, so the work that looks most finished now gets the shallowest critique.

Executiondesign reviewcritiqueDianne AlterThe Design Projectfidelitydesign leadershipAIprototyping
Helpful?

A brass balance scale weighing a single rough pencil figure sketch against a stack of glossy pixel-perfect interface screens. The sketch side sits heavily down, the polished screens float up weightless.

Dianne Alter posted a love letter to design reviews this week, and the part I keep rereading is not the part about AI.

Every Monday since 2020, The Design Project splits into breakout rooms. Everyone brings a project. Everyone gives feedback and gets it. Mandatory. And you can bring something you're stuck on or something you feel really confident about.

That second option is the whole thing. Stuck work gets fixed in a review. Confident work is where the team's bar actually moves, because that's the moment someone says the thing you didn't think of and you have to go re-open a file you'd already closed. Most teams only bring the stuck work, so their reviews function as a help desk and the bar never travels anywhere.

Then Dianne says the past year turned those Mondays into a place to share AI experiments with customers. Someone shows how they did something, and everyone else walks out with a new approach to try. That's not a design review anymore. That's a capability diffusion loop, and it's probably the most valuable meeting in her company.

The short version

The design review is not obsolete, but the signal it ran on is. For twenty years reviewers read fidelity as a proxy for conviction: grayscale meant early thinking, pixel-perfect meant someone had lived with the problem. AI broke that proxy, because a polished screen can be forty minutes old and carry no conviction at all. The predictable failure is that the work that looks most finished now gets the shallowest critique. The fix is to move the signal off the artifact and into the room: a stated confidence level before each presentation, the brief on screen before the pixels, and a critique aimed at the choice rather than the output. Options are free now, so "what did you kill and why" is the question that separates choosing from picking. This is the same argument as Prototype Before You Spec, arriving from the review side of the table.

The signal that broke

Here's what nobody has written down yet. For twenty years, reviewers read fidelity as a proxy for conviction, and it was a good proxy. A grayscale wireframe meant early thinking, so you challenged the concept. A pixel-perfect mock with real copy meant someone had already lived with the problem for two weeks, so you moved to spacing and states and edge cases. Nobody taught you that rule. You absorbed it.

That rule is dead. A polished screen with real photography and a working interaction can be forty minutes old and carry zero conviction behind it. The cost of looking finished collapsed. The cost of being right didn't move at all.

So reviews now fail in a specific, predictable way: the work that looks most done gets the shallowest critique. Everyone in the room is still running the old heuristic on a signal that no longer means anything.

The fix is embarrassingly cheap. Before anyone presents, they say one sentence about where they are and what they want. "Third direction, I don't believe it yet, attack the concept." Or, "I've been on this two weeks, I'm committed, find the failure modes." You're replacing a signal you used to read off the artifact with a signal the person states out loud. Ten seconds. It changes what the next twenty minutes are worth.

What to review when the artifact is free

Options are free now. Twelve directions is not twelve times the thinking. It's often zero times the thinking, and the reviewer can't tell the difference by looking.

So stop reviewing the output and start reviewing the choice. The question that separates the two is "what did you kill, and why." If someone can't name what they rejected and the reason it lost, they didn't choose a direction. They picked one. Picking is what a generator does. Choosing is the job.

This is the same move I make with evals. The useful signal is the wince, the moment a person looks at output and something is off before they can articulate why. A design review is the only place most teams still catch that reaction out loud, in a room, from someone other than the author.

Same logic applies to the brief. When generation was expensive, a weak brief showed up as a thin deck, and everyone could see it. Now a weak brief produces a beautiful, confident, wrong artifact that survives a whole review because it photographs well. If you only get one change out of this post, make it this one: put the brief on screen before the work.

The silo got more comfortable

Dianne opens with something I don't want to skip past. Solo designer at a startup, or surrounded by designers in corporate America, and both can be lonely when you're in your own silo.

AI made that silo nicer to live in. It answers instantly. It never says "I don't get it." It never says "I've watched three customers do this and they always go the other way." It will hand you a gorgeous answer to a bad question and it will do it agreeably, forever.

The friction that used to push a designer out of the silo and toward another human is mostly gone. Which means the pull has to come from somewhere else. It isn't a discipline problem, it's a missing ritual. I spend a lot of time arguing for killing the status meeting and every other ritual that outlived its job, and this is the other side of that: reviews are the ones you keep. Dianne's team made Monday mandatory in 2020, and the interesting part is that the same ritual now solves a problem that didn't exist when they built it.

Four things I'd put on the agenda

If you run one of these, or you're about to start one:

  1. A stated confidence level from whoever presents, before the screen share. One sentence.
  2. The brief and the constraint first, the pixels second.
  3. A "how I made this" slot. Two minutes, the actual method: the prompt, the tool, the eval, the thing that failed twice before it worked. This is how a team's capability stops being individual. TDP built a whole community around exactly this, which tells you what it's worth.
  4. Someone assigned to the confident work specifically. Rotate it. Their whole job that week is to take the piece nobody is worried about and try to break it.

I've been running product and design reviews for two decades, at Adobe, at Salesforce, and in every CPO seat I've held since, and the ones that mattered had almost nothing to do with the quality of the feedback in the room. They worked because they were the only recurring event where taste got calibrated across people who otherwise never saw each other's raw work. That was true before any of us had a model in the loop. It's more true now, because every designer on your team is quietly building a different private toolkit, and the bars are drifting apart faster than they used to.

Design is not the lost discipline right now. It's the one with the most leverage, and the teams that figure this out first will be the ones that kept a room where people look at each other's work and say the uncomfortable thing. More on where I think the practice goes next in the design operating model.

One thing to try this week: at your next review, ask everyone to state their confidence level before they present. Watch what happens to the conversation about the work that looked finished.

Credit where it's due: this whole line of thinking came out of Dianne Alter's post about The Design Project's Monday reviews. Go read hers.

Related answer: Why is a polished mockup no longer evidence of conviction?

Sources: Dianne Alter, "My love letter to design reviews and why AI will never remove them from my process", LinkedIn, 8 September 2026 · The Design Project, Miami · Design to Code Community, The Design Project.

Share this post

Also on Medium

Full archive →

Frequently asked

Why is polish no longer a signal of conviction in a design review?+

Because the cost of looking finished collapsed and the cost of being right did not move. For twenty years reviewers read fidelity as a proxy for how much thinking sat behind the work: grayscale meant early, pixel-perfect meant committed. A polished screen with real copy and a working interaction can now be forty minutes old. The proxy still feels true, which is what makes it dangerous.

How should a design review change now that AI can generate a polished screen in minutes?+

Replace the signal you used to read off the artifact with a signal the person states out loud. Before anyone presents, they say one sentence about where they are and what they want from the room: 'Third direction, I do not believe it yet, attack the concept,' or 'Two weeks in, I am committed, find the failure modes.' It takes ten seconds and it decides what the next twenty minutes are worth.

What should reviewers critique when generating options is free?+

The choice, not the output. Twelve directions is not twelve times the thinking, and often it is zero times the thinking, which a reviewer cannot detect by looking. The question that separates the two is what did you kill and why. If someone cannot name what they rejected and the reason it lost, they did not choose a direction, they picked one. Picking is what a generator does.

Why does a weak brief do more damage now than it used to?+

When generation was expensive, a weak brief showed up as a thin deck and everybody in the room could see it. Now a weak brief produces a beautiful, confident, wrong artifact that survives a full review because it photographs well. The single highest-value change to a review agenda is putting the brief and the constraint on screen before the pixels.

Should AI replace design reviews?+

No, and the argument runs the other way. An AI tool answers instantly, never says 'I do not get it,' and never says 'I have watched three customers do this and they always go the other way.' It will hand you a gorgeous answer to a bad question, agreeably, forever. The friction that used to push a designer out of the silo toward another human is mostly gone, so the pull has to come from a ritual instead.

What belongs on a design review agenda in 2026?+

Four things: a stated confidence level from the presenter before the screen share, the brief and the constraint before the pixels, a two-minute 'how I made this' slot covering the prompt, the tool, and what failed twice before it worked, and one person assigned each week to attack the work nobody is worried about. Rotate that last assignment.

THE SHORT ANSWER

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.