Most PM hiring loops in 2026 are optimized for a role that no longer exists. The loop was designed around 2018 skills: a case study, a prioritization exercise, a strategy framing, a behavioral round, a culture fit. Every candidate has been through this loop at least ten times and has been coached, by an LLM, a bootcamp, or a mentor, on exactly how to pass it. Signal-to-noise is so degraded you cannot tell a candidate who can do the work from one who memorized the shape of the answer. The fix is to stop testing the wrong thing.
Why each old round fails now
The case study tests packaging, not thinking: you learn whether they can make a pretty deck, not whether they can build a product. The prioritization exercise tests taste in frameworks, so you hire cultural alignment, not capability. The strategy round rewards the candidate most fluent at reasoning aloud, which is a real skill but not the one you need when the job is shipping. The behavioral round is fully exploited, since every candidate has an LLM-rehearsed answer to "tell me about a time you dealt with conflict." All four rounds test whether someone can behave like a PM. Now you need to know whether they can build.
The new loop
Round one, the builder task, a two-to-four-hour take-home. Send a real customer transcript, ask them to identify the opportunity, build a working prototype, write a 20-pair eval set, estimate cost per action, and write a one-page ship plan with an explicit rollback condition. Round two, a 60-minute review session where they walk you through it. The most revealing question is "what part of this are you least sure about?" Candidates who name their own weakest assumption with specificity are the ones you want. Round three, a 90-minute live pairing session: take a prompt from your production system, try to break it, improve it, run it against your eval set, watch the score change. This round cannot be faked. Prior practice shows in the first ten minutes. Round four, a 45-minute culture conversation with one question: tell me about a product decision you got wrong, what the signal was, and what you did about it.
Signals and the pushback
Weigh these heavily: they shipped something in the last 30 days even if small, they use Claude Code or equivalent as a daily tool and can say what they shipped with it last week, they talk about customers in the present tense, and they admit uncertainty with specificity. The loudest internal pushback is "we cannot hire this way because no one can pass." That is exactly the point. The current loop is easy because the skills it tests are widely practiced. The builder loop is hard because the skills it tests are scarce and the ones you actually need.
Pick one thing this week. If you have an open req, swap one round of your current loop for the four-hour builder task. Score it on: did it run, eval set quality, cost realism, rollback maturity. Compare what you learned to what the case study would have told you. One loop at a time, and within two quarters your hires are visibly different.