
I have made thousands of product decisions across twenty-some years, and for the first stretch of that career I learned from almost none of them. Not because I did not care. Because nothing was written down. By the time an outcome arrived, months later, my memory had quietly rewritten what I knew at decision time, what I expected, and how confident I was. Hindsight is a forger, and it forges in your favor.
Judgment is trainable. I have watched it improve in people I have managed, and slowly, embarrassingly, in myself. But it trains like a sport trains: reps plus a scorecard. Most PMs get the reps, hundreds of decisions a year, and skip the scorecard entirely. Experience accumulates. It does not compound.
The short version
Decision quality and outcome quality are different things, and conflating them (what Annie Duke calls resulting) is why orgs learn the wrong lessons from their own history. The training loop has four parts: a decision log with eight fields as the scorecard, written confidence percentages scored quarterly as the calibration habit, the premortem as the repeatable drill, and reversibility (one-way versus two-way doors) as the dial for how much process a decision deserves. This is the deep dive on decision quality from The Skill Stack, and the instrument is the decision log template. Speed of deciding is set by the cost of being wrong, not by temperament.
Resulting: the bug in how orgs learn
Annie Duke spent years as a professional poker player before writing Thinking in Bets, and the central idea transfers cleanly to product: a decision can be good and the outcome bad, because luck sits between them. Judging the decision by the outcome is what she calls resulting, and it is the default mode of every product org I have ever seen.
Resulting teaches backwards. The PM who shipped a well-reasoned bet that got run over by a market shift gets cautious. The PM who skipped discovery and got lucky gets bold, and promoted. Run that loop for a few years and the org has trained itself on noise.
The fix is not deep. It is keeping score on the part you control: the decision as it stood at the time. Which requires a record of the time, which is the log.
The decision log: eight fields
This is the training instrument. Mine has lived in plain text, a Notion table, and a spreadsheet at different jobs; the container does not matter. The fields do:
- Decision. One sentence. "We are pausing the export roadmap for two weeks to fix report latency."
- Date.
- Evidence at the time. Three to five bullets of what you actually knew. Not what you later learned. This field is the anti-forgery device.
- Confidence. A percentage. Not "fairly confident." A number you can be wrong about.
- What would change my mind. The observable thing that would flip this call. If you cannot fill this field, you have a belief, not a decision.
- Review date. When the outcome should be knowable. Put it on the calendar.
- Outcome. Filled in at review. What happened, plainly.
- Lesson. Filled in at review. About the decision process, not the outcome. "I weighted the loud account too heavily" is a lesson. "It worked" is not.
Two minutes per entry. Log significant decisions only: bets, kills, sequencing calls, hires, pricing moves. Five to ten a month is plenty. The template has the format ready to copy.
The unexpected benefit: the log improves decisions on the way in, not just the review on the way out. Writing "what would change my mind" before committing has stopped me from at least a few decisions that were really just momentum.
Calibration: the quarterly scoring
The confidence field is where the log becomes training rather than journaling.
Every quarter, pull the decisions whose review dates have passed. Bucket them by stated confidence. Across your 70% calls, roughly 70% should have gone the way you expected. Across your 90% calls, roughly 90%.
The first scoring is humbling for almost everyone. The common patterns: overconfident about timelines and adoption, underconfident about whether customers will pay, wildly overconfident in domains you know least (the less you know, the fewer counterexamples you can imagine). The specific shape of your miscalibration is the single most useful thing the log produces, because it is correctable.
And it changes how you talk in meetings. Once you have scored yourself for a few quarters, you stop saying "this will definitely work" because you know what your "definitely" historically converts at. You start saying "I am at 70 on this, and the thing that would move me is the next two enterprise calls." That sentence does more for your credibility than any deck, because almost nobody else in the room can say it.
The premortem: the repeatable drill
If calibration is the scorecard, Gary Klein's premortem is the rep. Before committing to a big bet, gather the team and announce: it is one year from now, this failed completely. Everyone writes the story of why, independently, for ten minutes. Then read them out.
Klein's research on prospective hindsight found that framing an outcome as certain materially improves people's ability to generate its causes. In practice the framing also does political work: it makes the skeptics safe. The engineer who would never say "I think this will fail" in a kickoff will happily write three vivid paragraphs about how it already did.
As a judgment drill, the premortem trains the exact muscle resulting atrophies: generating the failure hypothesis space before reality picks one. Run it on every one-way-door decision, log the top three failure stories in the evidence field, and at review time check which story, if any, came true. That comparison is a judgment rep no course can give you. The narrative craft of writing the premortem so it actually changes behavior is in storytelling is a PM core skill.
Bet sizing: doors and dials
Two more pieces complete the kit.
Reversibility. Sort decisions into one-way doors (hard to undo: pricing migrations, deprecations, public commitments, senior hires, data model choices) and two-way doors (cheap to undo: most feature bets, copy, defaults, experiments). The weighting rule: two-way doors get decided fast, by whoever is closest, with whatever evidence is on hand, because the cost of a wrong call is a quick reversal. One-way doors get the full kit: premortem, log entry, wider input, deliberately slower clock. Most orgs invert this, agonizing over reversible UI choices while waving an irreversible pricing change through on a single slide. Auditing your own last month of decisions against this rule is uncomfortable and useful.
Speed. When to decide fast versus slow is not a personality trait, it is a function of the cost of being wrong, which I unpacked in the cost of being wrong. Cheap-to-be-wrong plus reversible: decide today. Expensive-to-be-wrong plus irreversible: buy evidence first, and the premortem is the cheapest evidence available.
The 4-week practice plan
Week 1: start the log with five backfilled decisions. Pick five significant calls from the last quarter where you still remember the decision-time evidence. Backfilling teaches the format and gives the first review something to chew on.
Week 2: add confidence numbers. Log this week's decisions live, each with a percentage and a "what would change my mind." Notice how often the number is hard to commit to. That difficulty is the skill loading.
Week 3: run your first premortem. Pick the biggest bet currently in flight. Fifteen minutes with the team, silent writing, read-out, log the top three failure stories.
Week 4: review one aged decision. Take the oldest backfilled entry whose outcome is now knowable. Fill in outcome and lesson. Score the decision, not the result: given the evidence field, was it the right call? Write the lesson about process. That single review is your first real judgment rep with a scorecard, and the loop is now running.
The resource review, opinionated
Annie Duke, Thinking in Bets. The essential text. The resulting concept and the bet framing pay for the book in the first third. The later chapters on group decision processes are skimmable on first read.
Gary Klein's premortem piece in HBR. Four pages, the whole method. Read it before your week-three drill. Klein's longer work on naturalistic decision making is good but optional; the HBR piece is the payload.
Farnam Street. Worth skimming for the mental models inventory, with one warning: reading about models is not training. Use it as a reference, not a curriculum.
What to skip: pop-sci cognitive bias listicles and most "thinking better" content. Knowing fifty biases exists describes the problem without building the training loop. You do not debug overconfidence by reading about overconfidence. You debug it by writing 70% on a line, waiting a quarter, and getting scored. The log is the whole technology.
Pick one thing this week
Open a blank page tonight and backfill one decision: the biggest call you made last month, with the evidence you had then and the confidence you had then, as faithfully as memory allows. Set a review date. That is the first rep. The rest is just not stopping.
Sources: Annie Duke, Thinking in Bets, Gary Klein, Performing a Project Premortem, HBR.
Further reading
- The Decision Log Template, the instrument, ready to copy
- The Cost of Being Wrong, the dial for deciding fast versus slow
- The Skill Stack: What PMs and CPOs Must Learn Now, where decision quality sits in the twelve skills
- Storytelling Is a PM Core Skill, writing the premortem narrative so it changes behavior
Frequently asked
What is the difference between decision quality and outcome quality?+
Decision quality is whether the choice was good given the evidence available at the time. Outcome quality is what happened afterward, which includes luck. Annie Duke calls judging decisions by outcomes 'resulting,' and it is the main reason orgs learn the wrong lessons: good decisions get punished after bad luck, reckless ones get celebrated after good luck. You can only train judgment if you score the decision, not the result.
What fields belong in a PM decision log?+
Eight: the decision in one sentence, the date, the evidence available at the time, your confidence as a percentage, what would change your mind, a review date, the outcome (filled in later), and the lesson (filled in later). Two minutes per entry. The evidence-at-the-time field is the most important one, because it is the only defense against hindsight rewriting the story.
Why write down confidence as a percentage?+
Because 'probably' and 'likely' are unfalsifiable, and unfalsifiable predictions cannot train anything. Writing 70% creates a scoreable claim: across all your 70% calls, roughly 70% should come true. Scoring that quarterly reveals whether you are overconfident or underconfident, and in which domains. It also changes how you talk in meetings, because you start hearing the difference between 55% and 90% conviction in your own arguments.
What is a premortem and why does it train judgment?+
Gary Klein's exercise: before committing, assume the project has already failed and have everyone independently write the story of why. The prospective-hindsight framing surfaces risks that direct questioning misses, partly because it makes dissent safe. As a judgment drill, it forces you to generate the failure hypothesis space before reality picks one, which is exactly the muscle resulting never builds.
How should reversibility change how PMs decide?+
Sort every decision into one-way doors (hard to reverse: pricing migrations, public commitments, deprecations, key hires) and two-way doors (cheap to undo: most feature bets, copy, experiments). Two-way doors should be decided fast with whoever is closest to the work. One-way doors deserve the full kit: premortem, decision log entry, wider input, more evidence. Most orgs invert this and agonize over reversible things while waving irreversible ones through.
How long until a decision log shows results?+
The log is useful on day one as a thinking aid, because writing the evidence and confidence down improves the decision being logged. The compounding starts at the first quarterly review, when you score aged decisions against their recorded confidence. Most people find their first real calibration lesson within one quarter, usually some flavor of 'I am overconfident about timelines and underconfident about customer behavior.'

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn