The customer-value job did not change. A PM and a Product Builder are both trying to figure out which customer pain to solve next and prove they solved it. What changed is the medium they work in and how fast the loop runs. The cleanest way to see the difference is line by line.
Output, unit of work, cycle time
The old PM's core output was a document: a PRD, a spec, a one-pager, a roadmap slide. Engineering read it, asked questions, and built something slightly different. A Product Builder's output is a working artifact: a clickable prototype, an agent prompt, an eval rubric. There is nothing to interpret. The prototype either works or it does not.
The unit of work moved from the feature to the hypothesis. A feature is a thing you ship, scoped and estimated up front. A hypothesis is a thing you test. The Product Builder does not commit to building before the prototype tells them what to commit to.
Cycle time moved from weeks and months to hours and days. The old roadmap was built for weeks-long cycles because that was the speed of building. Cycles collapsed. The roadmap did not. That mismatch is the single biggest source of waste in old-shape orgs.
Why the change happened now
Building got cheap. A Product Builder using Claude Code can go from hypothesis to clickable artifact in an afternoon. When being wrong cost three months of engineering, you wrote a twelve-page spec to insure against it. When being wrong costs one afternoon, you build the thing and learn from the prototype. That drop in the cost of being wrong is the whole game.
The handoff changed with it. It used to be "here is what to build, please build it." Now it is "here is the thing that works, please make it survive a million users." Definition of done moved from a spec's subjective checklist to an eval that passes against a real test set with named slices. There is nothing left to argue about in the review.
What stays the same
Three things did not move. Customer judgment, knowing which pain to solve next, is still the hardest call in the building, and AI does not help much with it. Opinionated prioritization when the evidence is incomplete is still your call. And taste: two prototypes that score the same on the eval can be wildly different products, and the one a person wants to use is the one with taste.
If you lead a product org still optimized for the old column, run the three-team experiment. Pick three teams, give each PM a prototype-first mandate, kill the PRD requirement for a quarter, and replace the weekly status update with a weekly prototype review. Compare cycle time and ship rate against the rest of the org. The numbers end the debate.