How does a decision log train product judgment?

THE SHORT ANSWER

It forces you to record your confidence as a number and what would change your mind before the outcome arrives, then scores the decision quality separately from the outcome at a quarterly calibration review. That separation dismantles resulting, Annie Duke's term for judging a call purely by how it turned out. Over a quarter you learn where your stated confidence and your actual hit rate diverge, which tells you exactly where to trust your own gut.

Ask a product org for its decision log, the last ten significant decisions with the evidence that existed at the time and what happened after. Almost none can. The decisions live in Slack threads, meeting memories, and the heads of people who have since changed teams. So the org cannot learn from its own calls. It learns by anecdote, and anecdote is dominated by whoever tells the story loudest. Startups relitigate the same decision every six months. Large orgs promote people for outcomes that were mostly luck. Both failures have the same root: no record, no learning loop. The fix costs ten minutes a week.

The schema

One row per significant decision, eleven fields: an ID, the date, the decision in one line, the door type, the evidence present at the time, a confidence percentage, what would change your mind, a review date, and later the outcome, a quality verdict, and the lesson.

Two fields do most of the work. Confidence as a number forces you out of pretty sure into a falsifiable claim. You cannot calibrate pretty sure. You can calibrate 75 percent. What would change my mind is the pre-committed kill switch, written before the outcome arrives, which protects you from moving the goalposts later, something everyone does and nobody notices themselves doing. The door type, one-way or two-way, comes from Bryar and Carr's framing from Amazon. Two-way doors get a row and a fast call. One-way doors get a row plus a premortem.

The rule that trains judgment

In the moment, fill the first eight fields, five minutes. The hard rule: write the evidence field with only what you knew then, no retroactive polishing. If the evidence was thin, write evidence was thin. The outcome, verdict, and lesson stay empty until the review date. That gap is deliberate. Filling the verdict at decision time is just confidence restated. Filling it at review time, with the outcome known but scored separately, is where the resulting trap gets dismantled.

Calibration

Once a quarter, 45 minutes. Reread every row past its review date, fill in outcomes, score verdicts. Then bucket your confidence numbers: of everything you marked around 80 percent, what fraction actually worked out? The gap between your stated confidence and your hit rate is your calibration error, and it is almost never uniform. Most people are well calibrated in one domain and badly overconfident in another, usually the one they enjoy most. Knowing which is which changes how you weigh your own gut, which is the whole point.

For one-way doors, add a premortem based on Gary Klein's prospective hindsight: assume the decision failed twelve months out, write the failure story, list the three most plausible causes.

This week, download a log, reconstruct the last three significant decisions you remember, then log the next real one live with a review date. Put a recurring ten-minute block on Friday titled log and a 45-minute block on the last Friday of the quarter titled calibration. In a year you will have what almost no PM has: a scored record of your own judgment.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-07-31 · 3 min read