Field Notes
Methods from practice: evals, discovery, measurement, prioritization.
- Are prompt logs the new switch interview?For what customers want and where your product fails, prompt logs beat the call: their own words, the exact moment, and a sample of everyone. Interviews still win on why.
- Can you trust your product intuition?Less than you think. Intuition is pattern-matching on a world that may no longer exist, and AI rotates the patterns faster than any gut can recalibrate.
- Does jobs-to-be-done still work in the AI era?JTBD assumes the underlying job is stable. AI absorbs jobs faster than roadmaps react. The framework still names the job, it just stopped telling you how long the answer holds.
- What should you measure weekly versus quarterly for an AI product?AI features iterate ten times a week but outcomes take weeks to attribute. Measure direction weekly with leading indicators, outcomes on a slow cadence.
- How does a PM ship their first pull request with AI tools?The four-level PM PR ladder, PLANNING.md files in git, and the two non-negotiables that earn engineering trust before your first PR lands.
- How do you build a PM second brain from meeting recordings?PMs forget 90% of meeting context within a week. A four-step pipeline turns every recording into a searchable knowledge base that compounds over time.
- How do you do discovery when your customer is an AI agent?Agents don't answer interviews. Six methods replace the playbook: agent telemetry, failure-mode interviews with operators, capability audits, and more.
- How do you run a 15-minute sprint retro that improves things?Flip the retro ratio. AI generates a Sprint Health Report before the meeting from git, Jira, and Slack, so 15 minutes goes to decisions, not data gathering.
- How do you run your first product trio?A product trio is PM, designer, and tech lead making decisions as a unit. Two hours weekly: 50 minutes discovery together, then decide together with committed ownership.
- How do you run a weekly review that keeps you shipping?A three-phase weekly review in 30 minutes instead of 3 hours: gather the brief with AI, triage with five decision questions, communicate a standup opener.
- How do you train product judgment deliberately?Judgment trains like a sport: reps plus a scorecard. Use a decision log, written confidence scored quarterly, premortems, and one-way versus two-way doors.
- How do you write for machines, not just executives?Half your readers are agents now. Swap prose PRDs for eval-set-plus-brief, version prompts like code, and score every artifact with the cold-read test.
- How do you write OKRs that don't suck?Most OKRs are disguised task lists. Use the behavior change test: if the key result were hit, how would customer, user, or business behavior actually change?
- How does a decision log train product judgment?A decision log separates decision quality from outcome quality, records your confidence as a number, and calibrates your judgment against what actually happened.
- How do you avoid survivorship bias in product management?You only hear from the users who stayed. Talk to churned users, track non-events, run feature post-mortems, and put churn drivers in prioritization.
- How do you build a working prototype in 60 minutes?The exact five-phase workflow: frame the problem, describe to Claude Code, iterate UX, add real data, deploy and share. A testable app in one hour.
- How do you build a signal map in your first 30 days as a PM?Skip the coffee chats. In your first 30 days, map the seven places truth enters the building and audit whether each gets captured, synthesized, and routed.
- How do you measure the cost of being wrong for an AI feature?Speed is free now. Triage decisions by reversibility, not size. Ship cheap-to-undo calls instantly and slow down on the ones you cannot reverse.
- How do you run a customer interview that actually works?Show up with full signal context, run the 30-minute prototype interview format, and let AI synthesize the transcript in minutes. Depth over volume.
- How do you run continuous discovery?Agents ingest every call, ticket, and NPS response, extract signals, and surface a ranked opportunity brief every Monday. Then you prototype same day.
- How do you test a product assumption in a week?Map assumptions Monday, build a prototype that tests the riskiest ones, test with five customers, decide Friday. The prototype is the experiment.
- What is downside exposure and how do you score a feature on it?Downside exposure asks how fast a silent 20% quality drop would cost you customers. Score every feature on three axes: revenue through the workflow, cost of a wrong answer, and reversibility.
- What are direction metrics?Leading indicators measured on the cadence of the work itself. Seven signals that predict AI outcomes 4-8 weeks ahead, run on a two-layer system.
- What can no product framework teach you?Frameworks are training wheels. They teach the moves, but not judgment under ambiguity, which only comes from making real decisions and being wrong at scale.
- What does it mean to ship with observability?No feature leaves staging without the traces, metrics, and evals that tell you whether it works, before your first customer hits it. A seven-item contract.
- What does 'the eval is the spec' mean?The eval set replaces the PRD for AI features. 30 to 200 real input/output pairs define what good looks like, scored daily. The eval is the contract.
- What is an opportunity solution tree?An OST connects a business outcome to customer problems, candidate solutions, and experiments. Teresa Torres created it. AI populates it in minutes now.
- What is mob prototyping?One day a week, your PM, designer, and engineer build a working prototype together in one room. 24 person-hours instead of 41. Better output than solo.
- What is the impact loop?A four-beat operating rhythm that replaces sprints: Sense, Build, Measure, Amplify. It optimizes for responsiveness, not predictability. Same-day to eight days.
- When does vibe coding get more expensive than agentic engineering?Vibe coding starts cheap, then costs 3 to 10x more per feature past the crossover point. The CapEx versus OpEx cost curve, from Google's SDLC whitepaper.
- Which mental models break when building gets cheap?Four decision frameworks quietly price building as scarce and invert when AI makes it cheap: Opportunity Cost, Bottleneck, Sunk Cost, and Local vs Global.
- Why can't users describe what they need?Users know their problems deeply but prescribe solutions limited to what they know exists. Feature requests are symptoms, not specs. Show, don't tell.
- Why should you design the tournament instead of picking winners?Cheap prototyping did not kill product judgment. It moved the decision from picking the winner to seeding the bracket: which eight of fifty ideas get prototyped at all.
- Why do cheap prototypes sometimes kill good ideas?When building is nearly free, the demo replaces the argument. A rough prototype of a strong idea reads as a weak idea, and it dies in review.
- How do you write an eval rubric for an AI feature?Grade thirty real outputs on gut feel first, extract the rubric from your own disagreements, tie every dimension to a cost, and score binary, not 1-5.