The Rubric Is the Spec
Your taste currently lives in a recurring meeting, which caps quality at the number of things you can personally look at. A rubric extracted from real work is the version of that taste that executes when you are not in the room.
The short version
A design rubric is the spec, because it is the only form of taste that executes when you are not in the room. Write it by extraction, not invention: grade thirty real outputs on gut feel, find the phrases that repeat in your own reasoning, and keep only the dimensions where failing costs something you can name out loud. Score each dimension pass or fail with a severity attached, because a five-point scale lets two reviewers write 3 for different reasons and believe they agreed. Calibrate with two people on five outputs before you call it real. The template at the end takes one afternoon for the first block, and it is the download this handbook expects to travel furthest.
Your design taste lives in a meeting. Probably Thursday.
That is the constraint nobody names out loud. Quality in most companies is capped at the number of things one or two senior people can personally look at, and that ceiling was already the bottleneck before generation got cheap. Now a team can produce more work in a Wednesday than a review can absorb in a month.
You can respond by reviewing faster, which is the same job done worse. Or you can write the judgment down so it runs on every output instead of the ones that reached your calendar.
Why you extract a rubric instead of writing one
Almost every rubric I have seen fail was written from first principles, on a good afternoon, by someone smart. It came out with clarity, consistency, and delight on it. Then it scored every output a 4 out of 5 forever and quietly stopped being opened.
A rubric written that way describes a product nobody shipped. Its dimensions are the ones you would expect a design team to care about, which is exactly why they carry no information. They cannot separate this week's output from last week's.
Meanwhile the dimensions worth scoring are already in the sentences you say in review. It invented a number. It buried the answer. It hedged when it should have committed. It sounds like a robot apologizing. Those sentences are specific to your product, they are the actual ways your product is good and bad, and they are sitting in your notes right now.
So the order is inverted. Grade first, then find the pattern in your own grading.
Block 1: thirty outputs, no criteria
Pull thirty real outputs from the surface, or from a prototype of it. Real ones, not the examples somebody assembled to make a point in a deck. Then grade each one good, mixed, or bad, with one sentence of why.
No criteria yet. That part is deliberate, and it will feel wrong for the first ten.
Being inconsistent here is fine. The inconsistency is the data.
Two things to watch as you go, because they are what the exercise is for. Phrases that repeat in your why column are your candidate dimensions. And the places where you graded two similar outputs differently are disagreements with yourself, which mark precisely where the rubric will have to get sharp. Those rows are worth more than the thirty grades.
We wrote down the rubric dimensions we expected an AI feature to have. Then we graded thirty real outputs by gut, and the one that mattered most wasn't on the list. It was sitting in my own notes the whole time.
Block 2: the cost test
Now you have eight or nine candidate dimensions and every one of them feels important. Most of them are not.
For each candidate, name what failing it costs. Money, time, trust, harm, or rework. Invents a figure costs you a user acting on a false number and the trust that came with it. Buries the answer costs a few seconds of reading and mild annoyance. Tone feels off costs, when you sit with it honestly, nothing you can name.
If you cannot name the cost, cut the dimension. One line, and it does most of the work in this chapter.
Aim for three or four that each hurt rather than eight that measure vibes. Every dimension you keep is one somebody scores on every review from now until you retire it, so the bar to add one is high and should stay high. A rubric that takes twenty minutes per output to apply will be applied once.
Block 3: binary, with severity
Each surviving dimension gets a binary test. Passes when, fails when. Not a scale.
Scales lie. Given 1 to 5, two reviewers will both land on 3, one because the output was slightly vague and one because it was slightly wrong, and they will leave the room believing they agreed with each other. That is worse than no score, because now the disagreement is invisible and documented. Pass or fail forces them to say which, and the argument happens where it is useful.
Severity does the work the numeric scale was pretending to do. A blocker means do not ship, and one instance is enough. A major means fix before the next release. A minor means log it and watch the rate, and a minor that shows up in half your samples was never a minor.
Write the fails-when column first. It is more specific, and it forces a precision that the passes-when column can fake indefinitely.
This is the same machinery as The Eval Is the Spec, pointed at interface quality instead of model output. Product teams already accepted that a written eval beats a demo. Design has the identical argument available and mostly has not made it.
Calibrate before you believe it
Thirty minutes, and it is not optional.
Two people score the same five outputs independently, then compare dimension by dimension. For every disagreement, ask which of three things it is.
The rubric is vague. Both readings are defensible from the words on the page. Rewrite the test. This is the most common result on a first pass and it is a win, not a failure, because you just found the sentence that was going to cause six months of quiet inconsistency.
The rubric is measuring the wrong thing. The two scorers agree on the reading and both feel the resulting score is wrong. That dimension needs replacing, not rewording.
They disagree about quality, for real. Rare. Escalate it, make the call, and write it down. That call is the actual taste you have been trying to encode, and it only ever surfaces under this kind of pressure.
Repeat until two people land within one disagreement across five outputs. Then the rubric is usable by someone who did not write it, which was the entire objective.
Version it, then decide what runs without a person
A rubric is a spec, so it gets a version number, a changelog line, and an owner.
Two rules keep it honest. The rubric changes when the product changes, never when a score is inconvenient. Lowering a bar so a release passes is the failure this whole system exists to prevent, and it never announces itself. It happens on a Friday, with good intentions, in a thread. Second rule: when a dimension changes, label the scores from before the change with the old version, otherwise your quality trend is measuring your own edits.
Once it is calibrated and versioned, sort every dimension into three buckets.
Objective tests become automated checks and run on every build. Fabricated figures, missing citations, format violations, forbidden content. Cheap, and there is no reason for a human to ever look at them again.
Judgment dimensions become model-scored samples. A fixed set of outputs, scored against your written test, spot-checked weekly by a person.
And some dimensions stay with a person, in a named step, with a named owner and a slot on the calendar. Do not pretend this bucket is empty. Pretending it is empty is how products end up technically correct and joyless, and no eval will ever tell you that happened.
What the rubric is not
It is not a replacement for looking at the work. It is what makes looking at the work fast, because most of the argument already happened when the rubric was written.
It is also not a performance instrument. The moment scores attach to people instead of outputs, honest scoring stops and the whole thing turns political inside a month.
And it expires. A rubric that has not changed in a year is describing a product from a year ago. The rules for changing it are in the next chapter, along with who scores and what a disagreement actually means. The version I have found holds up across teams is in The Design Rubric Template, written to be usable by someone who never read this chapter.
This week: do Block 1 and nothing else. Thirty outputs, gut call, one sentence each. It takes an afternoon, and what makes that afternoon different from being handed somebody else's framework is that you find the dimensions in your own notes, in your own words, about your own product. Do Block 2 next week. The short version to send to someone who needs the method without the argument is How do you write a design rubric that scores work without you?
The first rule in Design Just Got Promoted was to write the constraint before anything gets made. This is that rule with a scoring sheet attached.
Chapter 12 of a series on the design operating model. Next: who scores, what a disagreement actually means, and when to change the rubric instead of the work.
Take the template
The design rubric template, flagship download
Frequently asked
What is a design rubric?+
A design rubric is a written, versioned set of quality dimensions with a pass and fail test for each one, used to score real outputs from a product surface. It is the form of design judgment that runs when the person who holds that judgment is not in the room. Unlike design principles, every dimension has a test somebody else can apply and a severity that says what happens when it fails. A good one has three or four dimensions, not eight.
How do you write a design rubric?+
By extraction, not invention. Grade thirty real outputs from the surface on gut feel first, with one sentence of reasoning each and no criteria written yet. The phrases that repeat in your reasoning column are your candidate dimensions, and the places you graded two similar outputs differently are where the rubric has to get precise. Then apply the cost test and write each surviving dimension as a binary check.
Should design rubric dimensions be scored 1 to 5 or pass and fail?+
Pass and fail. Given a five-point scale, two reviewers will both write 3 for completely different reasons and walk out believing they agreed. Given pass or fail they have to say which, and the disagreement surfaces where you need it. Severity carries the weight a numeric scale pretends to carry: a blocker means do not ship, a major means fix before the next release, a minor means log it and watch the rate.
How many dimensions should a design rubric have?+
Three or four that each hurt, rather than eight that measure vibes. Every dimension you keep is one somebody scores on every review forever, so the bar to add one is high. The filter is the cost test: name what failing the dimension actually costs in money, time, trust, harm, or rework. If you cannot name the cost, cut the dimension.
How do you calibrate a design rubric?+
Two people score the same five outputs independently, then compare dimension by dimension. Each disagreement is one of three things: the rubric is vague and both readings are defensible, the rubric is measuring the wrong thing and both scorers feel the result is wrong, or there is a real disagreement about quality. The first is common on a first pass and it is a win. The third is rare, and resolving it is the moment your practice's actual taste gets set.
When should you change a design rubric?+
When the product changes, not when a score is inconvenient. Lowering a bar so a release passes is the exact failure the system exists to prevent, and it happens quietly, on a Friday, with good intentions. Version the rubric like a spec, record what moved and which output or incident caused it, and label old scores with the version so your quality trend is not just measuring your own edits.
Related reading
Chapters and essays on the same thread, across both handbooks.
Design Just Got Promoted
The claim that design is lost contains a category error. Drawing got cheap. Deciding got more valuable, and there is more to decide than at any point in the last twenty years.
You Cannot Mock a Distribution
There is no final state to draw when the output is assembled at runtime. What you specify instead is a band, and everything inside it ships without your review.
Design's Kill List
A ritual belongs on the kill list when it produces a receipt rather than a decision, and the receipt only ever counted because it was expensive. Nine of those, each with its replacement.
Recovery and Trust Repair
The thirty seconds after your product is confidently wrong decide whether the person keeps using it. Almost nobody designs that moment.