
I argued in Every Artifact Is Somebody's Promotion Packet that the PM ladder ran on documents because documents were the evidence, and that AI is collapsing the evidence. That post names the problem. This one is the replacement, because a critique without a rebuild is just a complaint.
If the spec is no longer the currency of advancement, something has to be. Here is what counts as evidence at each level when the deliverable is gone.
The short version
When documents stop being the promotion currency, leveling has to run on judgment instead of output, which is harder to see and harder to game. The evidence at every level is the same three things: a shipped surface with a metric and your fingerprints on why it moved, an eval you own that caught a real failure, and a bet made with the reasoning written down before the outcome was known. What scales across levels is not the volume of that evidence but the size of the bet you are trusted to make without a net and the blast radius when you are wrong. L4 owns one surface and proves the loop works. L5 owns a workflow and sets the eval bar others inherit. L6 owns a pillar and is trusted with bets that are expensive to reverse. Principal sets the standard the whole pillar is judged against. The constraint most orgs hit is not the rubric, it is whether their calibration room can read a prototype-and-eval workflow well enough to tell a good bet from a lucky one.
The shift in one sentence
The old ladder measured how much you produced. The new one measures whether you were right, and how you knew.
That sentence is small and the consequences are not. "How much you produced" can be counted from the outside by anyone: pages, launches, projects. "Whether you were right and how you knew" requires someone who understands the work to look at a decision and judge the reasoning behind it. The evidence moved from artifacts you can stack to bets you have to evaluate. Everything below follows from that.
What counts as evidence at each level
These map to the four levels in the Product Builder job ladder. The JDs describe the scope. This is the evidence a calibration room should actually read.
L4, Product Builder
The evidence is that the loop works in your hands. You took one surface, shipped a working version, put it in front of real users, and moved a metric you can name. You wrote an eval before you built and it graded something real. The bet you were trusted with was small and reversible, and you did not need a net.
The tell that someone is actually L4 and not an L3 with a nice deck: they can show you a thing that shipped, a number that moved, and the eval that told them it was ready. Not a plan for those things. The things.
L5, Senior Product Builder
The evidence is that you set a bar other people now inherit. Your eval rubric became the one the team uses. Your prototype-first habit changed how the two builders next to you work. You owned a multi-step workflow, not a single surface, and the bet you were trusted with got bigger: something that touched other people's work and would have been annoying to unwind.
The tell: when you left the room, your standard stayed. An L5's fingerprints are on how other people do the work, not just on what they shipped themselves.
L6, Staff Product Builder
The evidence is a bet that was expensive to reverse and paid off, with the reasoning on record from before you knew. You owned a pillar. You made a call the org could not easily walk back, a pricing shape, an architecture direction, a thing to kill, and you were right for the reason you wrote down, not a reason you found afterward. You set the eval and observability standard the pillar runs on.
The tell: the company would feel the hole if you stopped making decisions. An L6 is trusted with irreversibility, and the only honest evidence for that trust is a track record of irreversible bets that went the way the person predicted.
L7, Principal Product Builder
The evidence is that you set the standard the pillar is judged against and shaped how the market sees the category. You are not making more L6 bets. You are defining the rules an L6 operates inside: what good looks like, what the company will and will not build, what craft bar the flagship holds. Your reasoning is quoted by people who do not report to you.
The tell: other people cite you as the reason they made a call. A Principal's judgment propagates through decisions they were never in the room for.
Why this is hard to run, and why that is the point
Look at what every one of those levels leans on: a decision, made under uncertainty, with the reasoning available to check. That is the load-bearing evidence, and it is hard to grade. A document can be skimmed in the ten minutes before calibration. A bet cannot. To level someone on a bet, a reviewer has to know enough to tell whether the person was right for the right reason or simply lucky, and luck and judgment look identical in the outcome column.
This is the real constraint, and it is not a rubric problem. It is a reviewer problem. Most product orgs have calibration rooms full of people who learned to grade documents, because documents were the currency for twenty years. Ask those same rooms to grade a prototype-and-eval workflow and a decision log, and half of them cannot read the evidence. That gap is why the director layer is under pressure: a manager who can only evaluate a PRD cannot evaluate the work that replaced it.
So the new ladder does not just need a new rubric. It needs reviewers who can read the new evidence, which means the leveling problem and the manager-repricing problem are the same problem wearing two hats. You cannot fix one without the other.
The upside is that the new evidence is harder to fake. You can pad a document. You cannot pad a metric that did not move, or a bet that was wrong for the reason you predicted it might be. The ladder gets more honest at exactly the moment it gets harder to administer, and honest-but-hard beats easy-but-fake, as long as you build the room that can read it.
The single best piece of evidence
If I could keep only one artifact for promotion decisions, it would not be a document or a prototype. It would be a decision log: a record of bets made under real uncertainty, each with the reasoning written down before the outcome was known.
That log separates judgment from luck, which is the only thing leveling is actually trying to measure. A person who was right five times with sound reasoning recorded in advance is a different bet than a person who was right five times and reconstructed the logic afterward. The old ladder could not tell them apart because it measured output. The new one can, if you keep the record.
Try this week
If you are a manager, rewrite the evidence column of your leveling rubric before the next calibration cycle. Replace "authored the PRD for X" with "shipped X, moved metric Y, and here is the bet and the reasoning from before we knew." One line per level. Then look honestly at your calibration room and ask whether it can read that evidence. If your reviewers can only grade documents, fix that before you promote another person against a currency the work has already left behind.
If you are an IC, start your decision log today. One entry: a call you are making this week under real uncertainty, the reasoning, and what you expect to happen. In six months you will have the one piece of evidence the new ladder is built to read, and almost nobody else will have it yet.
Frequently asked
What replaces documents as promotion evidence for PMs?+
Records of judgment under uncertainty instead of records of effort. At every level the evidence is a shipped surface with a metric and your fingerprints on why it moved, an eval you own that caught a real failure, and a bet made with the reasoning written down before the outcome was known. What scales across levels is the size of the bet you are trusted to make alone and the blast radius when you are wrong.
How do you level a PM without deliverables to point at?+
You measure the scope of the judgment, not the volume of the output. L4 owns one surface and proves the loop works. L5 owns a workflow and sets the eval bar others inherit. L6 owns a pillar and is trusted with irreversible bets. Principal sets the standard the pillar is judged against. The question at every step is the same: how big a decision can this person make without a net, and how do they reason when they are wrong?
Why is judgment harder to calibrate than documents?+
A document can be skimmed in the ten minutes before a calibration meeting. A bet requires knowing whether the person was right for the right reason or just lucky, which means a manager has to understand the work well enough to tell the difference. Most leveling systems were built to grade legible output. Grading judgment requires managers who can read a prototype-and-eval workflow, and that is the constraint most orgs hit first.
What is the single strongest piece of promotion evidence now?+
A decision you made under real uncertainty, with the reasoning recorded before you knew the outcome, that turned out right for the reason you predicted. That is the hardest thing to fake and the truest signal of the next level, because it separates judgment from luck. A decision log beats a document library.
Does this ladder assume everyone becomes an engineer?+
No. It assumes everyone ships working artifacts and owns evals, which is a different skill from writing production code. The Builder PM validates with a prototype and grades outputs with a rubric. Engineering still owns the systems that survive a million users. The ladder measures product judgment expressed through working things, not lines of code.
What should a manager do first to adopt this?+
Rewrite the evidence column of your leveling rubric before the next calibration cycle. Replace 'authored the PRD for X' with 'shipped X, moved metric Y, here is the bet and the reasoning.' Then check whether your calibration room can actually read that evidence. If your reviewers can only grade documents, fix that before you promote another person against a currency the work has already left.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn