The Review Protocol, Thirty Minutes
Not consensus. Not a redesign. The meeting produces three things:
The template
The Review Protocol, Thirty Minutes
A rubric review is not a design critique with a scoring sheet stapled on. It is a different meeting with a different output. Thirty minutes, four segments, three roles, and a rule for what a disagreement means. Run it weekly on one surface, not monthly on everything.
The output of the meeting
Not consensus. Not a redesign. The meeting produces three things:
- Scores, per dimension, per output, recorded.
- A list of disagreements, classified.
- One decision: ship, fix, or change the rubric.
If a review ends without a decision, it was a critique wearing a rubric's clothes.
The roles
Scorers, two of them. They score independently before the meeting. Not during. Scoring in the room produces agreement that is really just deference to whoever spoke first.
A facilitator. Runs the clock, reads out scores, classifies disagreements. Does not score. This is a real job for thirty minutes and doing it while scoring means doing neither.
The owner of the surface. Present, listening, not defending. The owner's job in this meeting is to hear the scores, not to explain them away. If the owner is also a scorer, you have lost the independence that makes the scores worth collecting.
Three people minimum. Five is the sensible maximum. More than five and it becomes a presentation.
Before the meeting, one hour total
- Facilitator pulls the sample. Ten outputs, real, from the last week. Include at least two the team would rather not look at, chosen by an obvious rule (lowest confidence, most edited, most reversed) so nobody is accused of stacking the set.
- Both scorers score all ten independently against the current rubric version. Fifteen minutes each if the rubric is good.
- Facilitator tabulates and marks disagreements. Does not resolve them.
If the scoring takes more than twenty minutes per scorer, the rubric is too long. Cut a dimension.
The thirty minutes
Minutes 0 to 5: The tape
Facilitator reads the aggregate. Pass rate per dimension, changes since last week, blockers found. No discussion. This segment exists so everyone starts from the same numbers rather than from their own impressions.
Minutes 5 to 20: Disagreements only
Skip everything both scorers agreed on. Agreement is not interesting and it eats the clock.
For each disagreement, the facilitator asks one question: which of the three is this?
A. The rubric is vague. Both readings are defensible. Action: rewrite the test, this week, by a named person. Do not rewrite it in the room, that produces wording nobody has slept on.
B. The rubric is wrong. Both scorers agree on the reading and both feel the result is wrong. Action: flag the dimension for replacement, and do not let the current version block a release on it.
C. Real disagreement about quality. Rare. Action: the surface owner's manager or the design lead makes the call, in the room, and it is written into the decision record. This is the moment the practice's actual taste gets set, so it does not get deferred to a follow-up thread where it will die.
Two minutes per disagreement, hard. If a disagreement needs longer, it is a category C and it needs a decision, not more discussion.
Minutes 20 to 27: The decision
Ship, fix, or change the rubric. Say it out loud, name the owner and the date.
Ship rules that hold across teams:
- Any blocker in the sample means do not ship, full stop, no exceptions negotiated in the room.
- Majors are counted, not debated. Over your stated threshold, do not ship.
- Minors are logged and the rate is watched. A minor that appears in half the sample is not a minor, it is a major nobody wanted to name.
Minutes 27 to 30: The one change
Pick the single highest-value change to the rubric or the product coming out of this review. One. Named owner, done before the next review.
A protocol that produces four actions produces zero.
What disagreement rates tell you
Track this. It is the health metric for the whole system.
| Disagreement rate | What it means | What to do |
|---|---|---|
| Near zero | Rubric is trivial, or scorers are anchoring on each other | Harder sample, or check they scored independently |
| Low but present | Healthy | Nothing |
| Around a third | Rubric is vague on at least one dimension | Find the dimension, rewrite the test |
| Above half | The rubric is not shared understanding, it is a document | Recalibrate from scratch, five outputs, together |
The failure people miss is the first row. Perfect agreement usually means the scorers talked before scoring, or the dimensions are so obvious they were never worth measuring.
Cadence
Weekly, one surface, rotating. Thirty minutes on the calendar with the same three people.
Monthly reviews do not work for generated products. Model behavior, prompts, and retrieval all drift faster than a month, and by the time a monthly review finds something, four weeks of output shipped on the old assumption.
Two ways this meeting dies
It becomes a critique. Somebody starts redesigning in the room and the clock goes. The facilitator's job is to say "that is a good idea, and it is not this meeting."
Scores start attaching to people. The moment a score is read as a judgment of the person who made the thing, honest scoring stops. Score outputs, never people, and never bring these numbers into a performance conversation. Do that once and the system never recovers.
The one thing to do this week
Run it once, badly. Pull ten outputs, get two people to score them independently, and spend fifteen minutes on the disagreements. The first run mostly teaches you that your rubric is vaguer than you thought, which is exactly the finding worth having.
From "Scoring Your Own Work", chapter 13 of The Design Operating Model. falkster.com/design/scoring-your-own-work