Extraction, not invention. That is the method, and skipping it is why most rubrics die quietly.
Almost every design rubric that fails was written from first principles on a good afternoon by someone smart. It came out with clarity, consistency, and delight on it, scored every output a 4 out of 5 forever, and stopped being opened by the second month. The dimensions were the ones you would expect any design team to care about, which is exactly why they carried no information about your product.
The dimensions worth scoring are already in the sentences you say in review. It invented a number. It buried the answer. It hedged when it should have committed. Those are specific, they are yours, and they are sitting in your notes.
Block 1: grade thirty outputs with no criteria
Pull thirty real outputs from the surface, or from a prototype of it. Real ones, not the set somebody assembled for a deck. Grade each good, mixed, or bad, with one sentence of why.
No criteria yet, deliberately. Being inconsistent here is fine, because the inconsistency is the data. Two things to watch: the phrases that repeat in your why column are your candidate dimensions, and the places you graded two similar outputs differently are disagreements with yourself, marking exactly where the rubric will have to get precise. Those rows are worth more than the grades.
Block 2: the cost test
Now you have eight or nine candidates and all of them feel important. Most are not.
For each, name what failing it costs. Money, time, trust, harm, or rework. Invents a figure costs a user acting on a false number and the trust attached to it. Buries the answer costs a few seconds and mild annoyance. Tone feels off costs, honestly, nothing you can name.
The rule: if you cannot name the cost, cut the dimension. Three or four that each hurt, not eight that measure vibes. Every dimension you keep is one somebody scores on every review from now until you retire it.
Block 3: binary, with severity
Each surviving dimension gets a pass and fail test, not a scale.
Scales lie. Given 1 to 5, two reviewers both land on 3 for different reasons and walk out believing they agreed, which is worse than no score because the disagreement is now invisible and documented. Pass or fail forces them to name which.
Severity carries the weight the number was pretending to carry. A blocker means do not ship and one instance is enough. A major means fix before the next release. A minor means log it and watch the rate, and a minor showing up in half your samples was never a minor.
Write the fails-when column first. It is more specific, and the passes-when column can fake precision indefinitely.
Calibrate before you believe it
Thirty minutes, two people, the same five outputs, scored independently. Then classify every disagreement as one of three things.
The rubric is vague: both readings are defensible from the words on the page, so rewrite the test. Most common on a first pass, and it is a win.
The rubric is wrong: both scorers agree on the reading and both feel the score is wrong, so the dimension needs replacing rather than rewording.
Real disagreement about quality: rare, and this is the one that matters. Escalate, make the call, write it down. That call is the taste you were trying to encode, and it only surfaces under this pressure.
Repeat until two people land within one disagreement across five outputs. Then it is usable by someone who did not write it, which was the point.
Then sort dimensions into what runs automatically, what a model scores against your written test, and what stays with a person on the calendar. Do not pretend the last bucket is empty. Pretending it is empty is how a product ends up technically correct and joyless, and no eval will tell you that happened.
The full worksheet is The Design Rubric Template, and the argument for why the rubric is the spec is in The Rubric Is the Spec. Do Block 1 this week and nothing else.