
Dianne Na Penn is Head of Product for AI Research and Labs at Anthropic. She joined in 2023 as the company's first technical product manager, back when there were roughly five product engineers, and she has had a hand in shipping every Claude model from Claude 2 to Mythos. When someone with that vantage point lists what they have learned about building on the frontier, I read it twice.
What follows is not a recap of her conversation on Lenny's Podcast. It is my reaction to her thirteen takeaways, as a founder building an eval-first product company around most of these ideas. Some I have argued for a year and it was useful to hear them confirmed from inside the lab. A few sharpened my thinking. I will say where I agree, where I would extend it, and where I would push.
The short version
Penn's throughline is that building on the frontier is a different craft than building normal software, and the tools have to change to match. Evals replace PRDs because a probabilistic system needs a runnable definition of success, not a prose description of intent. Capabilities arrive discontinuously, so you need evals as a radar to catch the jumps before they sit unclaimed as overhang. The craft shifts from sweating pixels to reading trajectories and diagnosing why a run failed. Managers who stop building lose their theory of what is possible. And through all of it, human judgment gets more valuable, which is why we will need product people more than ever. My take on all thirteen is below.
1. Evals are the new PRDs
This is the headline and the one I have written the most about. Penn's framing is clean: a PRD describes what to build, an eval defines what success looks like once you have built it. Her team writes 30 to 40 representative examples per major feature, each a prompt with a golden answer, and the eval suite is the primary artifact. PRDs still exist, mostly to align engineering, legal, and safety, but the eval is the core.
I have argued the same from the outside, that the PRD is dead and the prototype and eval are the spec, and I have published the exact template and how I build the rubric underneath it. The extension I would add: the eval does not just replace the PRD, it collapses the spec, the acceptance test, and the regression suite into one artifact, and it is the only one of the three a probabilistic system respects. Prose describing intended behavior is a wish. A test of actual behavior is a contract.
2. You need frontier products to feel frontier models
Penn's example is precise. Opus 4.5 would not have had its breakout moment without a vehicle like Claude Code, and Claude Code would not have accelerated without Opus 4.5. The team had felt the magic internally for months. The inflection came when model intelligence and product vehicle finally met.
This is the point I would attach to everything else. A capability with no product around it is a capability nobody can feel, including the team that built it. The product is the instrument that makes latent intelligence legible, and it is also how you discover the model is not as good as the benchmark implied. You do not learn whether a capability is real by reading the model card. You learn it by shipping something that depends on it and watching what breaks, which is the whole reason I build the outcome-to-prototype loop the way I do.
3. Sweat the tokens as much as you sweat the pixels
For twenty years product craft meant sweating the pixels, the interface micro-decisions that separate a tool people tolerate from one they love. Penn's point is that the craft now has a twin. Increasingly the work is reading transcripts, studying failed trajectories, and diagnosing whether a miss was a hallucination, an overconfident answer, or a wrong tool call. Each diagnosis routes to a different team and a different fix.
I love this because it is unglamorous and exactly right. The interface of an AI product is the trajectory, and you cannot design what you have not read. This is the same discipline I push in how I build eval rubrics: grade real outputs before you theorize, because the failure you can name is the only one you can fix. A team that never reads its own transcripts is tuning a system it cannot see.
4. AI capabilities arrive discontinuously, and evals are how you catch them
The scaling papers carry two kinds of graph, Penn notes: the smooth curve of loss going down, and the jagged graphs of emergent capability. A model can go from failing 1+1 to getting it right every time with no smooth transition. Her warning is exact: "unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know." The result is product and user overhang, value already sitting in the model that no product has claimed.
I have circled this under a different name. My dual-cadence measurement model exists because the fast loop of model change cannot be tracked on the slow clock of quarterly outcomes. Penn gives the sharper version. Evals are not a launch gate. They are a continuous capability radar, and the whole point is to catch a step-change the morning it lands rather than the quarter after a competitor shipped on it.
5. Golden Gate Claude was a hidden inflection point
In mid-2024 researchers dialed up a Golden Gate Bridge feature inside the model and Claude became obsessed, weaving the bridge into every answer, including a spaghetti recipe. A small team shipped a live consumer experience in 24 hours and reached only about 2,000 people. Penn still calls it "one of those hidden inflection points of finding our identity," because it proved Anthropic could ship experiences authentic to its own research.
This is my operating thesis happening inside the company that builds the models. A research artifact became a shipped product in a day, with a tiny team, and the act of shipping taught the company something about itself that no planning cycle would have surfaced. Identity is not decided in a strategy offsite. It is discovered by shipping something real and noticing how it feels. Most enterprises would have spent those 24 hours writing a brief about whether to build it.
6. AI writing is the current jagged edge, and it is being actively sharpened
Why is AI writing still so recognizable? Penn's answer is that the technology is jagged. The last push made models dramatically more agentic, which made writing the new rough edge, and there are now active efforts on both the product and research sides to make Claude write better, tone and character included.
The operational lesson is to know exactly where the edge runs for your domain, because betting a product on a still-jagged capability is how you ship something that dazzles in the demo and fails in the wild. The edge also moves, fast, so the teams that win build for where it will be in two model generations, not where it is today. But you only earn the right to that bet if you have the evals from takeaway four to tell you when the edge actually moved.
7. Anthropic's Labs team exists to make discontinuous 10x to 1,000x bets
Labs pulls on threads that are not on the core roadmap and asks what the 10x, 100x, or 1,000x version would be. That posture produced Claude Code, MCP, Skills, and Claude Design. The pods are tiny, sometimes a single engineer, because, in Penn's words, "really large teams pursuing very ambiguous, large ideas end up being slowed down." Bets that do not work get revisited a model generation or two later.
The part most orgs cannot copy is the people. You have to select for founder types who can watch a bet get turned off and go again, which only works with real kill discipline. I have written about the prototype graveyard for exactly this reason. The graveyard is not a failure, it is the cost of making bets big enough to matter. An org that chases moonshots but cannot kill them does not have a Labs team, it has a backlog of expensive orphans.
8. We will need PMs more than ever
This is the takeaway I would keep if I only kept one. Penn's definition: "the role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner, and doing the relentless work to do that, that to me is the core of a product person, and we actually need more of that." As AI makes building easier, the scarce human skill becomes judgment.
I just published a piece arguing that the PM's job was never to make decisions, it is to make sure decisions get made, and that this becomes the entire job once agents build faster than organizations can decide. Penn is describing the same gravitational shift from inside a frontier lab. When the cost of building drops toward zero, knowing what is worth building, and getting the organization to commit to it, is the work. That is a person, and for now it is a product person.
9. Managers who aren't hands-on building are losing their theory of mind
Penn's onboarding plan for tenured PMs is identical to the one for early-career hires: read transcripts, talk to users, understand what good looks like. She carves out one or two active workstreams on every model release to keep her own sense of how capabilities are moving. The line is blunt enough that I wrote it down: "if you're not building yourself, you're not gonna make it."
This is the same thing I have called the AI noise tax, the leaders who defend their distance from the work with talk of strategy while their sense of the possible quietly rots. The fix is not glamorous. Build something yourself, this week, with the current models, not to ship it but to keep your intuition calibrated to reality instead of last year's demo. A leader whose theory of mind is stale approves the wrong things confidently, which is the most expensive kind of wrong.
10. Protect your brain from AI brain rot
Form your point of view first, then use the model as a sparring partner. Penn's rule is to delegate fully only where the writing matters less than the thinking, like a standardized monthly business review, and to shift into the reviewer's seat, because who signs off matters more than who wrote it.
I treat this as a hard line. Write the first draft of your own thinking before you ask the model, because the value you add is the judgment that decides whether its answer is any good, and you cannot exercise judgment you have outsourced. Note how neatly this rhymes with takeaway nine and with my own argument that the PM is the accountable owner of the call. Who signs off is the decision, and the decision is the job. Use AI to extend your reach, not to replace the part of you that decides.
11. Use AI to raise your EQ, not just your IQ
This is the one I did not expect and liked the most. Penn built a custom Claude skill based on the book Crucial Conversations that coaches her before high-stakes moments, helping her go deeper faster, build trust, and be more direct, and she shares it with the other managers on her team.
Almost every conversation about AI at work is about raising IQ, doing the analysis faster, writing the doc quicker. Penn is pointing at the other axis. The hardest parts of a leadership job are the human ones, the difficult conversation, the piece of feedback you keep softening, and those are coachable too. This is a genuinely new use pattern to me, and a reminder that a skill is not only a way to automate a task. It can be a way to rehearse a moment you would otherwise walk into cold.
12. Alignment makes Claude more useful, not less
The common assumption is that safety is a tax on capability. Penn's experience is the opposite. Reducing sycophancy, the model's tendency to agree with whatever you put in front of it, produces real improvements in judgment and output. She has used a research version of Opus to pressure-test Claude pricing decisions specifically because it pushes back. Her line: a thinking partner does not just agree with you, it should add to you.
This lands hard for me because it is downside exposure seen from the model side. A model you can trust not to invent a number, and to tell you when you are wrong, is a model you can put on a high-stakes workflow where a fabricated number or a flattering yes costs you a customer. Alignment is not the thing standing between you and usefulness. It is the thing that unlocks the usefulness with real money attached, because the willingness to push back is the product.
13. Claude's coding dominance began as a relatively small training change
The most humbling one for anyone who plans in feature lists. In 2023, Penn says, the company was under 200 people and "nobody said Anthropic and coding in the same sentence." The team noticed people writing long-form code and decided to train Opus 3 to be better at it. A relatively small change to the training run won Anthropic its earliest developer enthusiasts, and the bet was strategic too, since recursive self-improvement requires models that write code and use tools well.
This is the whole argument for investing in the substrate over the surface. The highest-leverage move is rarely the biggest line item. It is the small change at the right layer that compounds through everything built on top of it. I organize my own engineering around exactly this belief, that the substrate, the toolkit and guardrails and eval harness, is where leverage lives and features are what fall out of it. The teams that win are not the ones with the longest feature list. They are the ones who found the small change that made everything downstream better at once.
The thread that ties them together
Read the thirteen together and one loop appears. You build a real product so you can feel what the model can do. You read the trajectories and measure with evals so you catch the jumps the moment they arrive. You keep your own hands on the work and your own brain in the loop so your judgment stays calibrated. And you spend that judgment forcing the decisions that turn a capability into a shipped outcome.
That is not a frontier lab secret. It is available to any team willing to trade the comfort of the PRD and the strategy deck for the discomfort of shipping, reading, measuring, and deciding in public. What Penn's vantage point confirms is that even at the place building the models, the edge is not the model. It is the craft of building the product that makes the model felt, and the judgment to know what is worth building next. That part is still ours.
Sources: Anthropic's first technical PM on token maxing, the jagged edge, and living in the future, with Dianne Penn (Lenny's Podcast); Anthropic's Dianne Penn Says the Eval Has Replaced the PRD (The New Stack); Golden Gate Claude (Anthropic).
Frequently asked
Who is Dianne Penn at Anthropic?+
Dianne Na Penn is Head of Product for AI Research and Labs at Anthropic. She joined in 2023 as the company's first technical product manager, when there were about five product engineers, and has helped ship every Claude model from Claude 2 onward. These are my reactions to her thirteen takeaways from Lenny's Podcast, not a transcript.
What does 'evals are the new PRDs' mean?+
A PRD describes what to build. An eval defines what success looks like once you have built it. Penn's team writes 30 to 40 representative examples per major feature, each a prompt paired with a golden answer, and treats that suite as the primary artifact of product work. PRDs still exist, mostly to align engineering, legal, and safety, but the eval is the core.
Why do AI capabilities arrive discontinuously, and why do evals matter for it?+
Scaling papers show a smooth loss curve but jagged emergent-capability graphs. A model can go from failing 1+1 to getting it right every time with no smooth transition. Penn's warning: 'unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know.' Without evals as a radar, the improvement sits unclaimed as product and user overhang.
Why does Dianne Penn say managers must stay hands-on?+
Her onboarding plan for tenured PMs is identical to the one for early-career hires: read transcripts, talk to users, learn what good looks like. She carves out one or two active building workstreams on every model release. Her line is blunt: 'if you're not building yourself, you're not gonna make it.' A manager who stops building loses an accurate theory of what the models can now do.
Does AI reduce the need for product managers?+
The opposite. Penn calls the core of a product person going into the details of what users are trying to accomplish and doing the relentless work to make it actionable, and says 'we actually need more of that.' As building gets easier, the scarce human skill becomes judgment, and judgment is a product skill.
Does alignment make Claude less capable?+
Penn's experience is the reverse. Reducing sycophancy, the tendency to agree with whatever you present, improves judgment and output. She uses a research version of Opus to pressure-test pricing decisions specifically because it pushes back. As she puts it, a thinking partner should not just agree with you, it should add to you.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn