# The Review Protocol, Thirty Minutes

> A rubric review is not a design critique with a scoring sheet stapled on. It is
> a different meeting with a different output. Thirty minutes, four segments,
> three roles, and a rule for what a disagreement means. Run it weekly on one
> surface, not monthly on everything.

---

## The output of the meeting

Not consensus. Not a redesign. The meeting produces three things:

1. Scores, per dimension, per output, recorded.
2. A list of disagreements, classified.
3. One decision: ship, fix, or change the rubric.

If a review ends without a decision, it was a critique wearing a rubric's
clothes.

---

## The roles

**Scorers, two of them.** They score independently before the meeting. Not
during. Scoring in the room produces agreement that is really just deference to
whoever spoke first.

**A facilitator.** Runs the clock, reads out scores, classifies disagreements.
Does not score. This is a real job for thirty minutes and doing it while scoring
means doing neither.

**The owner of the surface.** Present, listening, not defending. The owner's job
in this meeting is to hear the scores, not to explain them away. If the owner is
also a scorer, you have lost the independence that makes the scores worth
collecting.

Three people minimum. Five is the sensible maximum. More than five and it becomes
a presentation.

---

## Before the meeting, one hour total

- Facilitator pulls the sample. Ten outputs, real, from the last week. Include at
  least two the team would rather not look at, chosen by an obvious rule (lowest
  confidence, most edited, most reversed) so nobody is accused of stacking the
  set.
- Both scorers score all ten independently against the current rubric version.
  Fifteen minutes each if the rubric is good.
- Facilitator tabulates and marks disagreements. Does not resolve them.

If the scoring takes more than twenty minutes per scorer, the rubric is too
long. Cut a dimension.

---

## The thirty minutes

### Minutes 0 to 5: The tape

Facilitator reads the aggregate. Pass rate per dimension, changes since last
week, blockers found. No discussion. This segment exists so everyone starts from
the same numbers rather than from their own impressions.

### Minutes 5 to 20: Disagreements only

Skip everything both scorers agreed on. Agreement is not interesting and it eats
the clock.

For each disagreement, the facilitator asks one question: which of the three is
this?

**A. The rubric is vague.** Both readings are defensible. Action: rewrite the
test, this week, by a named person. Do not rewrite it in the room, that produces
wording nobody has slept on.

**B. The rubric is wrong.** Both scorers agree on the reading and both feel the
result is wrong. Action: flag the dimension for replacement, and do not let the
current version block a release on it.

**C. Real disagreement about quality.** Rare. Action: the surface owner's manager
or the design lead makes the call, in the room, and it is written into the
decision record. This is the moment the practice's actual taste gets set, so it
does not get deferred to a follow-up thread where it will die.

Two minutes per disagreement, hard. If a disagreement needs longer, it is a
category C and it needs a decision, not more discussion.

### Minutes 20 to 27: The decision

Ship, fix, or change the rubric. Say it out loud, name the owner and the date.

Ship rules that hold across teams:

- Any blocker in the sample means do not ship, full stop, no exceptions
  negotiated in the room.
- Majors are counted, not debated. Over your stated threshold, do not ship.
- Minors are logged and the rate is watched. A minor that appears in half the
  sample is not a minor, it is a major nobody wanted to name.

### Minutes 27 to 30: The one change

Pick the single highest-value change to the rubric or the product coming out of
this review. One. Named owner, done before the next review.

A protocol that produces four actions produces zero.

---

## What disagreement rates tell you

Track this. It is the health metric for the whole system.

| Disagreement rate | What it means | What to do |
|---|---|---|
| Near zero | Rubric is trivial, or scorers are anchoring on each other | Harder sample, or check they scored independently |
| Low but present | Healthy | Nothing |
| Around a third | Rubric is vague on at least one dimension | Find the dimension, rewrite the test |
| Above half | The rubric is not shared understanding, it is a document | Recalibrate from scratch, five outputs, together |

The failure people miss is the first row. Perfect agreement usually means the
scorers talked before scoring, or the dimensions are so obvious they were never
worth measuring.

---

## Cadence

Weekly, one surface, rotating. Thirty minutes on the calendar with the same three
people.

Monthly reviews do not work for generated products. Model behavior, prompts, and
retrieval all drift faster than a month, and by the time a monthly review finds
something, four weeks of output shipped on the old assumption.

---

## Two ways this meeting dies

**It becomes a critique.** Somebody starts redesigning in the room and the clock
goes. The facilitator's job is to say "that is a good idea, and it is not this
meeting."

**Scores start attaching to people.** The moment a score is read as a judgment of
the person who made the thing, honest scoring stops. Score outputs, never people,
and never bring these numbers into a performance conversation. Do that once and
the system never recovers.

---

## The one thing to do this week

Run it once, badly. Pull ten outputs, get two people to score them independently,
and spend fifteen minutes on the disagreements. The first run mostly teaches you
that your rubric is vaguer than you thought, which is exactly the finding worth
having.

---

_From "Scoring Your Own Work", chapter 13 of The Design Operating Model._
_falkster.com/design/scoring-your-own-work_
