Field Notes
Methods from practice: evals, discovery, measurement, prioritization.
- What is Teresa Torres's three-step approach to AI evals?Error analysis, choose the eval type, run the loop. What the Product Talk guide says, what to steal first, and the step that turns an eval into a gate.
- Do forward deployed engineers scale, or is it just consulting?Forward deployment costs like services and has to earn like software. When the math works, the buyer's lock-in objection, and the signal to stop sending engineers.
- Does zero data retention mean the vendor keeps nothing from your AI agents?Salesforce promised at Dreamforce 26 that your data does not go in the models. True for the model layer. Here is what that promise does not cover.
- How is a forward deployed engineer different from a sales engineer?One line separates them: a sales engineer's job ends when the customer signs and an FDE's job starts there. Confusing the two is behind most stalled AI pilots.
- What is a forward deployed engineer?The role Palantir invented and OpenAI, Anthropic, Sierra, and Google Cloud hire for in 2026. What it is, what it is not, and why it is a Product Builder in disguise.
- When can you actually enforce a rule on an AI agent?Spotify could gate file reads over 350 lines and not boilerplate writes. The three-question test that separates rules you can enforce from documentation.
- What is Marty Cagan's fresh definition of the product role?Cagan maps Benedict Evans's three product skills onto problem discovery, value risk, and viability risk. What it says, and what the mapping leaves out.
- Can you retrofit context for AI agents?No. Context is the only agent layer captured at write time. Control planes, routing, and governance can be added later. The reasoning behind a decision can't.
- Why did Miro sell for only 2.3x ARR?Miro is profitable with ~$600M ARR and sold at 2.3x. Compression explains the range, not the position. The mechanism is legibility.
- What did Salesforce's Q2 FY27 earnings say about agentic AI?Adjusted EPS matched consensus. GAAP jumped on an Anthropic gain. And Salesforce shipped its CRM inside Claude. The three numbers that matter.
- What are bottom-up evals, and why can't an LLM write them?Top-down evals restate the brief, so a model drafts them well. Bottom-up evals come from reading real outputs. Here is the conversion rule.
- Why do customers buy your product but not renew it?The reason a buyer approves you and the reason users keep coming back are different things. Most teams instrument only the first one.
- How do you design a screen that is generated at runtime?Specify the range, not the screen: best case, worst acceptable, and what the product refuses to do. Everything between the last two ships without review.
- How do you write a design rubric that scores work without you?Extract the rubric from thirty graded outputs, keep only dimensions with a named cost, score binary with severity, then calibrate on five outputs.
- What do you do in the thirty seconds after your AI is confidently wrong?Six steps in order: detect, admit in place, contain, correct visibly, name the prevention, adjust the confidence posture. Written before you need it.
- Does the IKEA vs Salesforce story prove AI automation fails?No. The viral story is quoted as proof AI fails, but both halves are stories of AI working, and the comparison misreads the tech, the 47%, and the sourcing.
- What are malleable loops in software?Dave Killeen's term for improvement moving through the software itself: the product self-checks, files scrubbed defect reports agent-to-agent, and ships.
- What is outcome-based pricing for AI agents and how does it work?Pay per resolved case, escalations free, tokens irrelevant. How Sierra's outcome pricing works mechanically, why it beats usage pricing, and where it breaks.
- Why are SaaS revenue multiples collapsing in the AI era?Airtable, Chegg, Stack Overflow, and Smartsheet repriced for two reasons. How to tell an AI repricing from a rate repricing, and which one you face.
- Why is Sierra AI growing so fast?Sierra hit $100M ARR in seven quarters and ~$200M by mid-2026. The mechanism is six operating choices, led by outcome pricing that ties revenue to landing.
- Why did Airtable sell for only 2.7x ARR?Airtable's ~$480M ARR grows over 20%, yet it sold at 2.7x revenue. AI didn't take the revenue, it took the option. Here is the mechanism and the fix.
- Are prompt logs the new switch interview?For what customers want and where your product fails, prompt logs beat the call: their own words, the exact moment, everyone. Interviews win on why.
- Can you trust your product intuition?Less than you think. Intuition is pattern-matching on a world that may no longer exist, and AI rotates the patterns faster than any gut can recalibrate.
- Does jobs-to-be-done still work in the AI era?JTBD assumes the underlying job is stable, and AI absorbs jobs faster than roadmaps react. It still names the job, it stopped saying how long it holds.
- What should you measure weekly versus quarterly for an AI product?AI features iterate ten times a week but outcomes take weeks to attribute. Measure direction weekly with leading indicators, outcomes on a slow cadence.
- How does a PM ship their first pull request with AI tools?The four-level PM PR ladder, PLANNING.md files in git, and the two non-negotiables that earn engineering trust before your first PR lands.
- How do you build a PM second brain from meeting recordings?PMs forget 90% of meeting context within a week. A four-step pipeline turns every recording into a searchable knowledge base that compounds over time.
- How do you do discovery when your customer is an AI agent?Agents don't answer interviews. Six methods replace the playbook: agent telemetry, failure-mode interviews with operators, capability audits, and more.
- How do you run a 15-minute sprint retro that improves things?Flip the retro ratio. AI generates a Sprint Health Report before the meeting from git, Jira, and Slack, so 15 minutes goes to decisions, not data gathering.
- How do you run your first product trio?A product trio is PM, designer, and tech lead deciding as a unit. Two hours weekly: 50 minutes discovery together, then decide with committed ownership.
- How do you run a weekly review that keeps you shipping?A three-phase weekly review in 30 minutes instead of 3 hours: gather the brief with AI, triage with five decision questions, communicate a standup opener.
- How do you train product judgment deliberately?Judgment trains like a sport: reps plus a scorecard. Use a decision log, written confidence scored quarterly, premortems, and one-way versus two-way doors.
- How do you write for machines, not just executives?Half your readers are agents now. Swap prose PRDs for eval-set-plus-brief, version prompts like code, and score every artifact with the cold-read test.
- How do you write OKRs that don't suck?Most OKRs are disguised task lists. Use the behavior change test: if the key result were hit, how would customer, user, or business behavior actually change?
- How does a decision log train product judgment?A decision log separates decision quality from outcome quality, records confidence as a number, and calibrates judgment against what happened.
- How do you avoid survivorship bias in product management?You only hear from the users who stayed. Talk to churned users, track non-events, run feature post-mortems, and put churn drivers in prioritization.
- How do you build a working prototype in 60 minutes?The exact five-phase workflow: frame the problem, describe to Claude Code, iterate UX, add real data, deploy and share. A testable app in one hour.
- How do you build a signal map in your first 30 days as a PM?Skip the coffee chats. In your first 30 days, map the seven places truth enters the building and audit whether each gets captured, synthesized, and routed.
- How do you measure the cost of being wrong for an AI feature?Speed is free now. Triage decisions by reversibility, not size. Ship cheap-to-undo calls instantly and slow down on the ones you cannot reverse.
- How do you run a customer interview that actually works?Show up with full signal context, run the 30-minute prototype interview format, and let AI synthesize the transcript in minutes. Depth over volume.
- How do you run continuous discovery?Agents ingest every call, ticket, and NPS response, extract signals, and surface a ranked opportunity brief every Monday. Then you prototype same day.
- How do you test a product assumption in a week?Map assumptions Monday, build a prototype that tests the riskiest ones, test with five customers, decide Friday. The prototype is the experiment.
- What is downside exposure and how do you score a feature on it?Downside exposure asks how fast a silent 20% quality drop would cost you customers. Score on three axes: revenue, cost of a wrong answer, reversibility.
- What are direction metrics?Leading indicators measured on the cadence of the work itself. Seven signals that predict AI outcomes 4-8 weeks ahead, run on a two-layer system.
- What can no product framework teach you?Frameworks are training wheels. They teach the moves, but not judgment under ambiguity, which only comes from making real decisions and being wrong at scale.
- What does it mean to ship with observability?No feature leaves staging without the traces, metrics, and evals that tell you whether it works, before your first customer hits it. A seven-item contract.
- What does 'the eval is the spec' mean?The eval set replaces the PRD for AI features. 30 to 200 real input/output pairs define what good looks like, scored daily. The eval is the contract.
- What is an opportunity solution tree?An OST connects a business outcome to customer problems, candidate solutions, and experiments. Teresa Torres created it. AI populates it in minutes now.
- What is mob prototyping?One day a week, your PM, designer, and engineer build a working prototype together in one room. 24 person-hours instead of 41. Better output than solo.
- What is the impact loop?A four-beat operating rhythm that replaces sprints: Sense, Build, Measure, Amplify. It optimizes for responsiveness, not predictability. Same-day to eight days.
- When does vibe coding get more expensive than agentic engineering?Vibe coding starts cheap, then costs 3 to 10x more per feature past the crossover point. The CapEx versus OpEx cost curve, from Google's SDLC whitepaper.
- Which mental models break when building gets cheap?Four decision frameworks quietly price building as scarce and invert when AI makes it cheap: Opportunity Cost, Bottleneck, Sunk Cost, and Local vs Global.
- Why can't users describe what they need?Users know their problems deeply but prescribe solutions limited to what they know exists. Feature requests are symptoms, not specs. Show, don't tell.
- Why should you design the tournament instead of picking winners?Cheap prototyping did not kill product judgment. It moved the decision from picking the winner to seeding the bracket: which eight of fifty get built.
- Why do cheap prototypes sometimes kill good ideas?When building is nearly free, the demo replaces the argument. A rough prototype of a strong idea reads as a weak idea, and it dies in review.
- How do you write an eval rubric for an AI feature?Grade thirty real outputs on gut feel first, extract the rubric from your own disagreements, tie every dimension to a cost, and score binary, not 1-5.
- Should you build your own data governance for AI agents?You will not out-govern Alation, Collibra, or Purview. Plug into the governance stack the enterprise already has, and design three things now.