# The Falkster Corpus

The product practice of Falk Gottlob, as one file your assistant can read.

Thirty years of building product, most recently four CPO seats and the
Falkster.AI build, compressed into the claims, the chapters, the agent
blueprints, the templates, and the answers. It is opinionated on purpose.
Where it disagrees with the consensus, the disagreement is the point.

Source: https://falkster.com/corpus · Built 2026-09-19 · Author: Falk Gottlob
Contents: 49 handbook chapters, 5 design chapters, 44 agent blueprints, 198 answers, 261 posts

## How to use this

Paste this file into your assistant's project knowledge (Claude Projects,
a ChatGPT project, a Cursor rule file, an AGENTS.md), then work normally.
The point is not to ask it about the corpus. The point is that when you ask
it to size a bet, write a brief, or decide what to kill, it answers the way
this practice answers instead of the way the average of the internet answers.

Every entry carries a canonical link. When something here matters to a
decision, follow the link and read the argument. A summary is enough to act
on and not enough to disagree with.

## Attribution

Written by Falk Gottlob. Free to use for your own work and your team's.
When it shows up in something public, cite it as: Falk Gottlob, falkster.com,
with the canonical link. Not licensed for republication, resale, or model
training.

---

## The five claims

Everything here is downstream of these. Each links to the body of work arguing it.

### Enterprise AI Agents

Agent count is a vanity metric. What matters is how many deployed agents still complete production work after ninety days, and what each successful outcome costs.

Evidence: https://falkster.com/answers/topics/enterprise-ai-agents

### AI Product Management

AI did not make product management harder. It collapsed the cost of the parts PMs were trained to be good at, and left the parts nobody trained for.

Evidence: https://falkster.com/answers/topics/ai-product-management

### SaaS to AI Business Models

Software priced per seat is priced against a labor cost that AI removes. The pricing model has to move to the outcome, and the margin structure moves with it.

Evidence: https://falkster.com/answers/topics/ai-business-models

### Building a Company in Public

A founder writing while building has information nobody else has. The value is in publishing the decisions before the outcome is known, including the ones that turn out wrong.

Evidence: https://falkster.com/answers/topics/building-a-company

### Product Leadership

The product org chart was calibrated to a build cost that collapsed. Leadership work now is deciding which coordination roles stop being necessary, and saying so out loud.

Evidence: https://falkster.com/answers/topics/product-leadership

## The handbook (49 chapters)

Each chapter's own extractable summary. The full text is at the canonical link and in falkster-corpus-full.md.

### Why This Exists

Category: Foundation
Canonical: https://falkster.com/handbook/manifesto

The PM role is splitting into two tracks. Traditional PMs write specs and manage process. Product Builders use AI to collapse the distance between customer insight and working product. Builders prototype in hours using tools like Claude Code, validate with real customers the same week, and ship faster because they're learning continuously instead of planning quarterly. Three practices made the biggest difference in my own work: talking to customers every week (not quarterly), building clickable prototypes before writing any spec, and using AI agents for the mechanical work (monitoring dashboards, drafting reports, summarizing calls) so the actual PM job gets more time. Falkster.com is where I document what's working at the intersection of product management and AI, organized around the four areas I keep returning to: the Product Operating Model, Continuous Discovery, Outcome Orientation, and AI-Native Execution.

### The AI Product Operating Model

Category: Foundation
Canonical: https://falkster.com/handbook/product-operating-model

The AI product operating model has three phases. The pre-AI foundation (Marty Cagan's empowered teams, Teresa Torres's continuous discovery, the PM-designer-engineer trio) was correct and its core principles carry forward: customer obsession, small empowered teams, outcome measurement. What AI breaks: the PRD is dead because a working prototype is faster to build than a spec to write, the trio is becoming a quartet or a duo as role boundaries blur, discovery compresses from four weeks to four days when agents synthesize signals and generate prototypes automatically, and execution overhead (tickets, dashboards, status updates, release notes) drops from 40% of PM time to near zero. What the new model looks like: Monday starts with AI-synthesized insights instead of manual report-pulling, Wednesday produces working clickable prototypes instead of wireframes, Friday ships the first iteration instead of planning the next sprint. The biggest shift: Friday went from "plan what to build" to "ship what you built." That's not a tweak; that's a different operating rhythm.

### Kill the Roadmap

Category: Foundation
Canonical: https://falkster.com/handbook/kill-the-roadmap

The roadmap is the most expensive lie in product management. It freezes a plan based on last month's signals, rewards conviction theater, and creates a re-plan tax so high that teams execute against plans they know are wrong. I stopped publishing roadmaps 14 months ago and replaced them with a one-page live bet portfolio, updated every Monday. Three sections: active bets (5-7 things you are testing with hypothesis, signal, kill condition, and decision date), recently killed (what you stopped and what you learned), and standing queue (fewer than 20 things waiting for a trigger). The bet portfolio is not a planning failure. It is honest about uncertainty in a market that updates daily. The hardest part is emotional: trading the false comfort of claimed certainty for the real authority that comes from knowing what is actually happening right now.

### Old PM vs Product Builder, The Ledger

Category: Foundation
Canonical: https://falkster.com/handbook/old-pm-vs-product-builder

Product management got rewritten because the cost of being wrong collapsed. The old PM wrote specs to insure against expensive engineering bets, ran rituals to translate between functions, and shipped on a quarterly cadence. The Product Builder ships prototypes in an afternoon, writes evals as the contract, and runs on a daily eval-driven cadence. The unit of work moved from document to working artifact. The cycle time moved from weeks to days. The accountability moved from scope-and-timeline to outcome-and-cost-per-request. This ledger lays it out, line by line, so a CEO, CFO, or CPO can read it once and price the gap between their current org and where the market is going. If your product org is still optimized for the left column, you are paying for both jobs and getting neither.

### Continuous Discovery

Category: Discovery
Canonical: https://falkster.com/handbook/continuous-discovery-autopilot

Continuous discovery used to mean two customer interviews a week and a monthly synthesis session. AI changes the math entirely. Agents now ingest every sales call, support ticket, NPS response, and app review your company generates, extract signals automatically, and surface a ranked opportunity brief every Monday morning. When a signal emerges, you prototype in hours using AI coding tools and put something working in front of the customer the same day. The cycle that used to take four to eight weeks now runs same-day when the signal is clear, and inside a week when the problem needs a human in the room. Teresa Torres was right about the habit of continuous discovery. AI removed the excuse that you don't have bandwidth to do it.

### Kill the Status Meeting

Category: Leadership
Canonical: https://falkster.com/handbook/kill-the-status-meeting

The status meeting exists because nobody trusts the dashboard. Fix the dashboard once and the meeting disappears. Build one URL per product: six live strips covering product health, adoption, customer signal, active bets, cost, and incidents, all updated automatically. The meeting that was 45 minutes of prep plus 45 minutes of verbal repetition becomes a 20-minute optional walk of the page, or disappears entirely. At Smartcat, killing the status meeting and replacing it with a live product page reclaimed about 6 hours a week of calendar time per PM. The only meetings that survive are the ones where a human decision gets made.

### Engineering Builds the Substrate, Not Features

Category: Foundation
Canonical: https://falkster.com/handbook/substrate-first-engineering

In an AI-native org, engineering's highest-leverage work is not features. It is the substrate: the infrastructure and toolkit that lets PMs and designers ship working software into production safely. Four pieces make it up. Scaffolded environments, so a builder gets a safe, running sandbox in minutes instead of a week. Guardrails, so no build can touch or spend more than it should. An eval harness, so every build is scored against a bar before it graduates. Isolated deploys, so a prototype reaches a real customer without ever touching the real system. When the substrate is good, a hundred people can build and only the survivors reach production. When it is missing, every prototype is a risk and engineering becomes the queue everything waits in. The role does not shrink. It moves from writing features to owning the ground everyone else builds on.

### Build Your First Opportunity Solution Tree

Category: Discovery
Canonical: https://falkster.com/handbook/your-first-ost

An Opportunity Solution Tree connects a business outcome to customer problems, candidate solutions, and experiments. Teresa Torres created the framework. The structure has not changed. What has changed is speed: AI can populate the opportunity layer from hundreds of support tickets, sales calls, and NPS responses in minutes, where interviews used to take weeks. The new one-week loop is: Monday, review the AI-generated opportunity brief; Tuesday, build a working prototype for the top opportunity; Wednesday and Thursday, show it to five customers; Friday, decide and update the tree. At Smartcat I went from customer signal to shipped feature in four weeks using this loop. The old model took three to four months.

### The Interview Guide That Actually Works

Category: Discovery
Canonical: https://falkster.com/handbook/interview-guide

Customer interview technique has not changed, but everything around it has. AI agents prepare you before the call by pulling the customer's support history, usage patterns, and NPS score so you walk in already past the surface. During the call, real-time transcription frees you to listen fully instead of splitting attention with note-taking. After the call, a synthesis agent compares the transcript against hundreds of other data points in minutes. The prototype interview format, 30 minutes instead of 45, confirms an agent-identified signal, goes deep on the customer's workaround, and shows a working prototype to get a concrete reaction. Three to five interviews with AI prep and synthesis outproduce ten interviews done the old way.

### The Assumption Testing Playbook

Category: Discovery
Canonical: https://falkster.com/handbook/assumption-testing

Assumption testing is the practice of identifying the riskiest bets underneath a feature idea and using a prototype to test them in days, not weeks. The prototype is the test. Instead of designing separate experiments for desirability, usability, and viability, you build one working thing that tests all three at once. The one-week cycle: map assumptions Monday morning, build the prototype Monday afternoon, test with five customers Tuesday and Wednesday, synthesize Thursday, decide Friday. Red assumptions (the ones that kill the project if wrong) go first. A failed test is a win. It is the fastest way to kill a bad idea before engineering touches it. At Smartcat, one prototype test saved eight weeks of building the wrong recommendation system and led directly to a shipped feature that moved activation by 28%.

### Continuous Listening: Every Customer, Every Day

Category: Discovery
Canonical: https://falkster.com/handbook/continuous-listening

Continuous listening is a daily pipeline that ingests every customer signal source (support tickets, call transcripts, NPS surveys, churn reasons, product rage clicks), synthesizes overnight by theme, and surfaces the top clusters every morning in a five-minute digest. Weekly 1:1 interviews with Teresa Torres-style continuous discovery are still essential, but they now serve a different purpose: understanding the why behind a cluster you already detected, not finding the cluster in the first place. A PM with Claude Code can wire the v1 pipeline in two weeks. The old model, 5 calls a week as the only signal channel, gave you 250 data points per year from your most cooperative customers. The pipeline gives you thousands per day from everyone, including the customers about to churn who never take your calls.

### Prototype Before You Spec

Category: Execution
Canonical: https://falkster.com/handbook/instant-prototyping

Prototype before you spec. The PM who walks into a meeting with a working prototype beats the PM with slides every time. The 2-hour prototyping method: frame the problem in three sentences (who, what they're trying to do, how you'll know it's solved), build the core interaction in Claude Code or Cursor using experience-first descriptions, test immediately with one real person without explaining anything, then decide to kill or iterate. Vibe coding is not writing production code. It's translating a customer problem into a working artifact using clear English and conversational iteration. If you can write a clear email, you can do this. The barrier is artificial. Pick something you're genuinely unsure about and block two hours this week.

### The Impact Loop

Category: Execution
Canonical: https://falkster.com/handbook/impact-loop

The Impact Loop is a four-beat operating rhythm: Sense (know what is happening), Build (make a working response instead of a plan for one), Measure (quantify what actually changed), Amplify (scale what works and kill what does not). It replaces sprints because sprints optimize for predictability and the Impact Loop optimizes for responsiveness. The loop runs continuously, not on a fixed cadence. AI agents handle the sensing layer automatically, surfacing a two-minute daily brief. Prototyping takes hours, not weeks. Measurement is automatic and daily. A full loop from customer signal to validated, profitable change took eight days in the Smartcat example here. Compare that to the thirty-plus-day waterfall most teams run without noticing.

### The Eval Is The Spec

Category: Execution
Canonical: https://falkster.com/handbook/the-eval-is-the-spec

The eval set replaces the PRD as the primary specification artifact for AI features. An eval set is 30 to 200 real input/output pairs that define what good looks like, built from actual production inputs (support tickets, customer messages, usage logs), labeled with the correct output, spiked with 10 to 20 adversarial examples the system should refuse or handle carefully, and scored daily against a rubric (exact match, semantic similarity, LLM-as-judge, or human review, depending on the slice). Engineering builds against the eval set. Done means the score went up, not that someone approved a spec. Feature reviews shrink from 40 minutes of opinions to 20 minutes of score diffs and regression analysis. The first eval set takes an afternoon. By the third one, you have a template and wonder why you were ever writing PRDs.

### Building Evals: Error Analysis, Eval Types, the Loop, the Gate

Category: Execution
Canonical: https://falkster.com/handbook/building-evals

Four steps. One, error analysis by hand: read fifty outputs, write down the two or three mistakes that recur, rank them by what they cost a customer. Two, match each error to an eval type: code assertion for anything deterministic, golden dataset for one-right-answer questions, LLM-as-a-judge for semantic judgment, customer feedback signals for real usage, always cheapest first as a filter for the expensive one. Three, the loop: fixed inputs, baseline, one change, compare across every eval, log it. Four, the gate: for each action, decide whether it is reversible, and let the eval's pass history buy autonomy only on the reversible side. Correctness is the product team's definition, never the vendor's. Credit for steps one through three goes to Torres. The fourth is mine, and it is the step that makes the first three matter for an agent.

### Ship With Observability or Don't Ship

Category: Execution
Canonical: https://falkster.com/handbook/ship-with-observability

No feature leaves staging without the traces, metrics, and evals that will tell you whether it's working, before your first customer hits it. The instrumentation contract is a one-page document written before any code: success metric (the one number proving the feature works), leading indicators (early signals predicting whether the success metric will land), cost meter (real-time cost per successful action), eval set (named set with a pass threshold), trace points (which actions get logged), dashboard URL (exists before launch even if empty), and kill condition (metric below threshold for time period means the feature is reviewed for deprecation). If any of the seven is missing, the feature is code complete, not done. Observability and evals are the same loop from two angles. Watch both on the same page. A feature without observability is a feature you shipped on faith.

### The Deprecation Playbook

Category: Execution
Canonical: https://falkster.com/handbook/the-deprecation-playbook

Deprecate on signal, not politics. Every feature ships with an explicit kill condition written at launch: if weekly active usage stays below 2% of MAU for 8 weeks, the feature is reviewed for deprecation. That condition turns deprecation from a political debate into the execution of a decision already made. The three tiers: soft sunset (hidden from defaults but still accessible), hard sunset with migration (feature removed by a date with proactive customer comms), and immediate kill (broken or dangerous features only). The killer move on customer comms is individual notice to the specific users who used the feature in the last 90 days. And the cultural move that makes it all work: publicly celebrate what gets killed. A page showing features killed in the last 12 months, what the data showed, and what you learned reframes deprecation from failure to discipline.

### Incident Response Is a PM Ritual

Category: Execution
Canonical: https://falkster.com/handbook/incident-response

Incidents are the cheapest discovery your company already does. Every Sev-2 or worse exposes a latent product assumption that was wrong, a missing eval, and a mis-scoped workflow. The engineering post-mortem covers the infra fix; the PM post-mortem covers the product lesson and they are different documents that both need to happen. The PM-owned post-mortem has six sections: user-facing description, wrong assumption, eval gap, what customers said during the incident, product-level fix, and updated eval set. It is written within 48 hours, shared at the weekly team review, and the eval update is committed alongside it. Customer comms during the live incident are also the PM's job: acknowledge fast, estimate honestly, send the post-incident note explaining what changed. After three months of this practice, incident patterns become the leading indicator of product health that NPS misses by quarters.

### Agent-to-Agent Dispatch

Category: Execution
Canonical: https://falkster.com/handbook/agent-to-agent-dispatch

Agent-to-agent dispatch is what happens when you stop putting a ticket between the customer signal and the build. A listening agent sits on your calls, tickets, and churn surveys and extracts the outcome a customer is reaching for. Instead of writing that up for a human to triage, it hands the brief straight to a prototyping agent, which builds a working prototype the same day. By the time you read the morning digest, the prototype is already attached. You review the extraction and the prototype, then decide: show it to the customer, iterate, or kill it. No backlog, no planning meeting, no handoff. The PM stops initiating work and starts editing it. The control that keeps this sane is not the ticket queue, it is the eval bar: most dispatched prototypes die, and the survivors graduate to engineering hardening.

### PM AI Agent Fleet, Mapped to the 7-Stage Operating System

Category: AI Agents
Canonical: https://falkster.com/handbook/ai-agent-army

The PM AI agent fleet is a set of autonomous agents that cover every repeated decision and output a product manager makes, mapped to the seven stages of the PM Operating System: Sense, Discover, Decide, Build, Ship, Measure, Amplify. Each agent runs on a schedule, pulls from connected data sources via [Model Context Protocol (MCP)](https://modelcontextprotocol.io/), and delivers reports where you already work. Setup takes about an hour. Deploy them incrementally starting with Red Flag Detection. By month two the whole fleet is running. This page is the live index, auto-generated from the blog, and it updates every time I publish a new agent post.

### When Not to Use AI

Category: AI Agents
Canonical: https://falkster.com/handbook/when-not-to-use-ai

Before wrapping any surface in an LLM, walk a five-step decision tree in order and stop at the first yes: can a rule do it, can a query do it, can a form do it, can a heuristic and lookup do it, and only then use a model. The most powerful pattern in 2026 is the hybrid: use a rule for 80% of cases, escalate to a model for the 20% that actually need it. That hybrid cuts costs in half with no quality drop. The cost math is real: a 1,500-token prompt plus 500-token completion at 100 actions per user per day across 50,000 users is 1.5 million dollars a month for a feature a regex could do for free. The PM who can name why a model is necessary (rather than just convenient) is the one who earns credibility when the model genuinely matters.

### Gross Margin Is Your Job Now

Category: AI Agents
Canonical: https://falkster.com/handbook/gross-margin-is-your-job-now

Gross margin is now a PM decision, not a finance one. AI-first SaaS runs 55 to 70 percent gross margin against traditional SaaS at 78 to 85 percent, and the gap is entirely driven by product choices: model routing, prompt hygiene, caching, and early-exit logic. Three levers every PM can pull without becoming an ML engineer. First, build a routing table and send classification, extraction, and format conversion tasks to models 10x cheaper than your flagship. Second, cut production prompts ruthlessly: most are 3 to 10x longer than needed and cutting a prompt from 3,000 tokens to 600 typically saves 80 percent on that surface with no eval regression. Third, turn on prompt caching (50 to 90 percent cost reduction on repeated prefixes) and design early-exit rules for the top 20 percent of inputs by volume. Watch one metric: cost per successful action by surface. If you don't have it, finding out why not is your week.

### Pricing for AI Products

Category: Foundation
Canonical: https://falkster.com/handbook/pricing-for-ai-products

Per-seat pricing is broken for AI products because cost-to-serve now scales with usage, not with license count. Power users at a $50 seat can cost $200 to serve. Companies still on per-seat pricing in 2026 run gross margins 40 points below those on hybrid or outcome-based models. The four models that work: hybrid (base fee plus usage), outcome-based (pay per resolved ticket or successful action), tiered consumption (buy a bucket upfront), and pure usage (pay for what you use). The hardest decision is picking the right value unit: customers must be able to tell you what 100 units will do for them before signing. If they can't, you will churn them on their first big bill. Migrate without losing the book by grandfathering existing customers, selling hybrid to all new accounts, and re-pricing expansion on the new model only.

### Prompt Ops

Category: AI Agents
Canonical: https://falkster.com/handbook/prompt-ops

Prompt Ops means treating prompts with the same lifecycle as production code: version control, pull request review, automated evals on every change, staged rollout (1%, 10%, 100%), monitoring, and one-click rollback. The five-piece stack takes half a day to set up for one prompt. The PM owns the prompt (it is the spec, encoding what the system does, who it's for, and what tone it uses); engineering owns the wiring. Most teams break because prompts live in three places at once, edits are untracked, there is no testing layer, and rollback takes hours during a production outage. Start by moving one prompt into your repo this week, wiring your code to load from the file, and adding an eval set into CI. Build from there.

### The Living Changelog

Category: AI Agents
Canonical: https://falkster.com/handbook/the-living-changelog

The Living Changelog is a continuous eval replay against production. You run the same eval set every day against the live system and treat any delta beyond noise as an incident, even before a customer complains. Model vendors change behavior without telling you: snapshots get rerouted, safety training updates silently, quantization changes roll out without a release note. The system has three parts: a replay set (50-300 production examples), a daily automated run that stores scores with timestamps, and a drift alarm set at 2x the noise floor. When the alarm fires, a six-step runbook covers confirmation, vendor check, version pinning, fix decision, customer communication, and eval set update. Build the system for one surface this week and you will catch your first silent vendor regression before any customer does.

### Trust, Safety, and the Guardrail as a Product Decision

Category: AI Agents
Canonical: https://falkster.com/handbook/trust-and-safety

Every guardrail in an AI product is a product decision: what you refuse, what you warn on, what you silently log, what you allow with a disclaimer. Outsourcing that to legal produces a product that is both annoying and unsafe. The guardrail tier system has four levels: Tier 0 (hard block, list under 15 items or you are over-blocking), Tier 1 (soft warning, product proceeds with a flag), Tier 2 (logged only, product proceeds normally), Tier 3 (allowed, the default for most inputs). The most common failure is the overly cautious assistant that refuses legitimate requests and churns customers silently. Track refusal rate as a primary product metric alongside conversion. The real top three risks in most AI products are: the product states something false, the product reveals data it should not, and the product takes an irreversible action incorrectly. Address those before the legal team's list.

### Your Weekly Playbook

Category: Leadership
Canonical: https://falkster.com/handbook/pm-as-team

The week is the rhythm. The loop is a day. Monday is a strategic reset: read the overnight signal brief, review outcomes, set one focus. Friday is reflection: show working prototypes, tune the agent fleet, capture what you learned. In between, you run the outcome-to-prototype loop as many times as clear signals warrant, and each run collapses into a single day. A listening agent surfaces the outcome a customer is reaching for in the morning, a prototyping agent turns it into something clickable by midday, a real customer touches it in the afternoon, and you decide, kill, iterate, or graduate, by end of day. The slower, high-judgment work, deep interviews about hard problems and trio brainstorming, is the weekly layer that the fast loop does not replace. Most of what used to fill your calendar, standups, sprint planning, backlog grooming, status meetings, quarterly planning, is gone. What replaces it is customer evidence, working prototypes, and outcome data, produced in days instead of quarters.

### Strategy From Signals, Not Slides

Category: Leadership
Canonical: https://falkster.com/handbook/strategy-from-signals

The annual strategy deck is a memorial to one Wednesday in February. By the time it is published it is already a fossil, and the re-plan cost is so high that teams execute against assumptions everyone knows are stale. What I run instead: a one-page living strategy doc with two sections. Section one is 5-10 beliefs about the world, each tagged with confidence level and date last reviewed. Section two is, for each belief, the specific signal that would falsify it. Updated every Monday in a 20-minute solo review. Most weeks the note to leadership says "nothing material changed." Occasionally it says "belief X shifted, here is what that means for our bets." The annual offsite does not set strategy anymore; it reviews the year's evolution and identifies the least-confident beliefs. It would be easy to read this as a slower strategy deck. It is a different artifact, one that admits uncertainty instead of performing it.

### The Anti-Backlog

Category: Leadership
Canonical: https://falkster.com/handbook/the-anti-backlog

The backlog is a graveyard pretending to be an inventory. A 400-item Jira backlog is not a plan: it is institutional amnesia in tabular format. The anti-backlog replaces it with three things: a live queue capped at two weeks of work, signal-fed from customer signal clusters, eval regressions, incident action items, and cost spikes only; a hypothesis library for half-formed ideas written as hypotheses not features; and a kill list for decisions you have already made not to do something. Migration from a large backlog takes four weeks. Week 1: archive everything older than six months. Week 2: sort the rest by signal evidence. Week 3: stop adding to the old backlog. Week 4: delete it. The discomfort is highest in week 4. It passes. What replaces it is a team that ships in response to signal.

### The Builder PM 30/60/90

Category: Foundation
Canonical: https://falkster.com/handbook/builder-pm-30-60-90

The Builder PM 30/60/90 is the structured shift from traditional PM to product builder inside the job you already have. Days 1 to 30 are invisible infrastructure: Claude Code, a prompt repo, one automated agent, a personal signal dashboard, and three fewer recurring meetings. Days 31 to 60 ship one visible artifact that replaces old practice (an eval-as-spec, a live product page, or a 60-minute prototype). Days 61 to 90 make the change structural by documenting it, converting a peer, shifting one metric of success, and publicly killing one old artifact. You will feel like a fraud at day 1, exposed at day 20, threatened at day 45, and vindicated at day 70. Each is on schedule.

### Hiring the Builder PM

Category: Leadership
Canonical: https://falkster.com/handbook/hiring-the-builder-pm

Hiring the Builder PM means testing the actual skill: can this person ship a working prototype in four hours? The old PM hiring loop tests presentation, frameworks, and behavioral polish, all skills that coaching and LLMs have made easy to fake. The new loop has four rounds: a builder task (real customer transcript turned into a working prototype with an eval set and cost estimate), a review session (walk through what you built and defend the trade-offs), a pairing session (iterate on a production prompt against an eval set live), and one culture question about a decision you got wrong. Fewer candidates pass. The ones who do are visibly better.

### PM-as-Editor: Managing a Fleet of Agents

Category: AI Agents
Canonical: https://falkster.com/handbook/pm-as-editor

PM-as-Editor is the skill you need once your agent fleet is running. Most PMs either trust agent output blindly (broken telephone) or rewrite everything from scratch (worse than doing the work themselves). The skill that scales is editing: reading agent output the way a senior PM reads a team member's PRD, cutting what doesn't serve the purpose, shipping the 80% version instead of perfecting toward 100%. The operational framework is a four-tier trust ladder: Tier 1 ships without review, Tier 2 after a one-minute sanity check, Tier 3 after a real edit, Tier 4 the agent assists but you author. Every edit you make is training data for the prompt. Spend 60 seconds after each edit writing what you changed and why, then update the prompt. Do this weekly and the fleet gets sharper faster than any competitor can replicate.

### The PM Agent Stack: Open-Source Tools Mapped to PM Work

Category: AI Agents
Canonical: https://falkster.com/handbook/pm-agent-stack

The destination for every product organization is one AI brain with read access to every system the company runs. Slack, email, calendar, meetings, documents, source code, dashboards, CRM, design files. All of it. Not parts. All. Full stop.

### The Cannibalization Decision Framework

Category: Leadership
Canonical: https://falkster.com/handbook/cannibalization-decision-framework

The hardest CPO call during the AI inflection isn't whether to ship an agent product. It's whether to keep selling the SaaS product that pays for it. Four diagnostic questions route the decision. Can the legacy architecture support the successor's quality bar? Is the legacy customer base the right ICP for the successor? Can the company afford the gross margin trough (it compresses from 78-82% to 58-65% during the ramp)? Is the buyer the same person? Those answers route to one of three operating modes: sunset (replace the legacy on a published date), refresh (rebuild with AI capabilities under the same pricing), or split (run both product lines under one company). The political coalition takes 90 days to assemble across CEO, CFO, CRO, CCO, and yourself. The seven-decision sequence must run in order: sunset date, reorg, comp rewrite, board pre-sell, customer migration in cohort waves, one tracked metric, and post-sunset reorg planned six months out. Out-of-order kills the transition.

### Dual Transformation: Running Two Clocks

Category: Leadership
Canonical: https://falkster.com/handbook/dual-transformation

Dual transformation is two cadences, three talent categories, six CEO scoreboard numbers, and three rituals inside one product organization. The legacy SaaS runs on a six-week cycle with quarterly OKRs and seat-based forecasting. The agent-native successor runs on a one-week cycle with daily eval reviews, weekly outcome cohorts, and monthly pricing experiments. The talent rule is hard fences: a maintenance team (5 to 15 percent of engineering), a successor team (60 to 75 percent), and bridge engineers (5 to 10 percent) who own data, auth, identity, and infrastructure. The six CEO numbers are legacy revenue, NRR, and gross margin alongside successor outcome volume, gross margin per outcome, and percent of legacy revenue migrated. The three rituals are the Friday Read-Out (15 minutes, both teams, surface mid-week conflicts), the Sunset Drumbeat (monthly memo, migration progress), and the Quarterly Reality Check (honest review of cadences, talent allocation, and sunset-date adherence). Trying to run both products on one clock kills the new product first.

### Pricing Migration: The 18-Month Quarterly Playbook

Category: Leadership
Canonical: https://falkster.com/handbook/pricing-migration-18-month-playbook

This is the six-quarter operating playbook for migrating from per-seat to outcome-based pricing. Quarter 1: internal alignment, four mandatory conversations (CFO on the gross margin trough, CRO on comp plan, board on the trough narrative, lead customer on the pilot), and the contract template drafted. Quarter 2: hybrid pricing live for new customers, comp plan rewritten with legacy at 50% historical and successor at 150% ACV, lead customer becomes a public reference. Quarters 3 through 5: three-wave customer migration ending with the long-tail announcement and self-service tooling. Quarter 6: final migration push and post-sunset reorg. The gross margin trough bottoms at 58 to 65% in months 10 to 12 and recovers to 70 to 75% by month 24. Pre-sell the curve to the board in quarter zero. The five non-negotiable contract terms are unit definition, dispute window, arbitration mechanism, committed minimum, and price ceiling. The comp asymmetry is the lever: hold the 50% number even when the CRO pushes back.

### Direction Metrics for AI-Native Velocity

Category: AI Craft
Canonical: https://falkster.com/handbook/direction-metrics

Direction metrics are leading indicators measured on the cadence of the work itself. For AI-native and agent products, outcome metrics like NRR or CSAT lag 4-12 weeks behind your changes. With teams iterating 10-20 times per week, outcomes cannot drive day-to-day decisions. Direction metrics close that gap. The seven leading indicators that predict outcomes 4-8 weeks ahead: eval pass rate, agent quality score, iteration count, design coherence, customer escalation rate, dispute rate on outcome billing, and latency at p95/p99. Run two measurement layers: the seven leading indicators reviewed daily and weekly, and outcome metrics (NRR, CSAT, NPS, expansion) reviewed monthly and quarterly. Prevent gaming by auditing correlation quarterly. If an indicator drops below 0.5 correlation with the outcome it predicts, replace it.

### The Investor and Board Narrative for AI-Era SaaS

Category: Leadership
Canonical: https://falkster.com/handbook/investor-and-board-narrative

The AI-era SaaS investor narrative has four components: gross margin compression (accepting 55-70% GM during ramp versus the 75-85% of mature SaaS in exchange for higher absolute revenue), NRR redefinition (classical seat-based NRR breaks for outcome pricing, so you report two numbers: committed minimum NRR and outcome volume growth), the new rule of 40 (rule of 35 during the trough, recovering to rule of 40 as unit economics mature), and the right comp set (transition-peer companies like Sierra, Intercom Fin, and HubSpot AI, not pure-SaaS comps that make the trough look catastrophic). The enabling move is co-authoring the narrative with the CFO across three working sessions before any board meeting, so both of you are presenting the same numbers with the same framing. The board needs to be retrained on the new metrics over two quarters, not surprised by them.

### The Pricing-Tier Sunset Playbook

Category: Leadership
Canonical: https://falkster.com/handbook/pricing-tier-sunset-playbook

A pricing-tier sunset is not feature deprecation. It is the end of a revenue line, and it carries different stakes, different politics, and different communication. The four-stage structure spans 18 to 24 months: internal preparation (months 0 to 6), strategic account migration via Wave 1 individual calls (months 6 to 9), mid-market and long-tail migration (months 9 to 18), and sunset day plus post-sunset reorg (months 18 to 24). The comp asymmetry is the lever: legacy at 50% historical comp, successor at 150% equivalent ACV. Five patterns kill most sunsets before they close: Wave 1 calls delegated to account managers, the comp plan watered down by the CRO, customer cohorts treated uniformly when edge cases need separate tracks, the dispute mechanism untested at scale, and sales reps quietly negotiating extensions. Audit these five first.

### The CPO 30/60/90

Category: Leadership
Canonical: https://falkster.com/handbook/cpo-30-60-90

The CPO 30/60/90 is built around four audits and three artifacts. Days 1 to 30: run the reality audit, the money audit, the quality audit, and the decision audit, while logging every commitment you make in a trust ledger. Days 31 to 60: instrument what the audits exposed, build your coalition map, and ship one visible decision that signals the new bar. Days 61 to 90: write the kill list, place your first two or three bets, and deliver a day-90 readout structured as an SCQA memo to the exec team and board. Do not reorg, do not rewrite strategy in week 2, and do not spend the 90 days collecting opinions you cannot falsify. The templates for the ledger, the interview guide, and the readout are in [the CPO First-90 Kit](/blog/cpo-first-90-kit).

### The PM 30/60/90

Category: Foundation
Canonical: https://falkster.com/handbook/pm-30-60-90

The PM 30/60/90 for a new job runs in three moves. Before day 1 and through days 1 to 30: set up your personal stack before you start, run decision archaeology on the last 90 days of your product area, build a signal map instead of a stakeholder map, and write a manager contract in week one. Days 31 to 60: ship one compounding artifact, a call synthesis, a live dashboard, an eval set, or a debate-ending prototype, and use it to earn your seat in decision rooms. Days 61 to 90: take ownership of one bet with a falsifiable success measure, kill or renegotiate one inherited commitment, and write a day-90 note to your manager that sets the scoreboard for the next year. Templates for the contract, the signal map, and the day-90 note are in [the PM First-90 Kit](/blog/pm-first-90-kit).

### The Skill Stack: What PMs and CPOs Must Learn Now

Category: Leadership
Canonical: https://falkster.com/handbook/skill-stack

The skill stack has five layers. Judgment: problem selection, decision quality, taste. Expression: storytelling, writing for machines, prototyping. Systems: evals, agent orchestration, signal architecture. Economics: unit economics, pricing. Leadership: coalition building and killing things well. IC PMs live in expression and systems while building judgment; CPOs are paid for judgment, economics, and leadership. The two highest-ROI skills to start with are evals (rarest, most employable) and storytelling (multiplies everything else). Every skill below has a self-test; run them all, and spend the next two quarters on your two worst scores, not your two favorites.

### The Forward Deployed Engineer Is a Product Builder in Disguise

Category: Foundation
Canonical: https://falkster.com/handbook/fde-is-the-product-builder-in-disguise

A forward deployed engineer is an engineer employed by a vendor, embedded inside a customer, and accountable for the outcome the software produces there. Palantir popularized the role around 2008 and ran four vintages of it, each adding responsibility without shedding the last: platform stability, then data integration, then custom solutions, then enablement. The thread that survived every vintage is customer accountability. The role exploded in 2026 because code got cheap and outcome pricing arrived, and both make the person who guarantees the result the scarcest seat in the company. That seat requires the same six skills this handbook attributes to the Product Builder. The FDE got there from services. The Product Builder got there from product. They are the same job with different badges.

### FDE vs Sales Engineer: The Job Starts Where the Contract Is Signed

Category: Foundation
Canonical: https://falkster.com/handbook/fde-vs-sales-engineer

A sales engineer's accountability ends at the signature. A forward deployed engineer's accountability begins there. Every other difference, on-site time, integration depth, ownership of code, follows from that one. Solutions architects design the fit and hand off the build. Consultants deliver recommendations and leave. Professional services implements a fixed spec for a fee. Developer relations scales across thousands of users through docs. The FDE goes deep into one customer's production system and is measured on whether the outcome the customer bought actually shows up. The test for any of these titles is not what they are called. It is who is on the hook after go-live.

### Becoming an FDE: The Transition From PM, Engineer, or SE

Category: Execution
Canonical: https://falkster.com/handbook/fde-transition

The FDE role needs three things: production engineering, customer proximity, and accountability for the outcome. PMs have the second and third and need to close the engineering gap, which a coding agent has made a quarter's work rather than a career's. Engineers have the first and often the discipline behind the third, and need to move their desk to the customer's building. Sales engineers have the first two and need to stay past the signature, which sounds trivial and is the hardest of the three because everything in their compensation and calendar pulls the other way. In every case the 90-day plan is one real deployment, owned end to end, measured on the customer's result, with someone senior shadowing the first one. The role is senior by nature. Only 3 percent of open FDE roles in mid-2026 were early-career.

### What Has to Change in the Product Before FDEs Can Work

Category: Execution
Canonical: https://falkster.com/handbook/product-changes-for-fde

Five things have to be true in the product before the first FDE lands. There must be a configuration boundary, a deliberate line between what an FDE can change per customer without code and what needs the product team, and most customer-specific work has to land on the configuration side. There must be a per-customer sandbox with guardrails, so the FDE builds outcomes instead of environments. There must be an eval harness, because the executable definition of the outcome is the contract the FDE works to and the thing outcome pricing bills against. There must be an observation record per customer, because an FDE who cannot explain why the number moved cannot own it. And there must be a feedback path from deployment to roadmap that has an owner, or the loop that made Palantir's model compound never closes. All five are product work. None of them are the FDE's job to build.

### The Deployment-to-Product Loop: How FDE Work Becomes Roadmap

Category: Execution
Canonical: https://falkster.com/handbook/fde-to-product-loop

The deployment-to-product loop has two halves with two different owners. FDEs build the fast, overfit version for one customer and are measured on that customer's outcome. A product team with the explicit job of generalizing takes what was built, reads the observation records across customers, and turns the third occurrence of the same pattern into a feature or a configuration. The kill discipline is the same as for prototypes: one customer's workaround dies, three customers' identical workaround gets productized. The PM sits at the junction as editor, deciding what earns generalization and writing the eval the generalized version has to pass. When both halves have owners, each deployment makes the next one cheaper. When only the first half does, the company is a consulting firm with a software company's burn.

### FDE Economics: Margins, Outcome Pricing, and When to Stop

Category: Leadership
Canonical: https://falkster.com/handbook/fde-economics

Forward deployment compresses gross margin in the short run, and the model only earns it back if the deployment-to-product loop closes and the pricing puts the vendor's revenue behind the outcome. Outcome pricing is the reason the model is spreading: billing per result requires someone to guarantee the result, and that is the FDE. The buyer's objection is lock-in and dependency, and it is legitimate, but the fix is contractual: the observation record and the configuration belong to the customer and port with them. The signal to stop sending engineers is when the product's configuration surface has absorbed enough that the customer, or their own AI engineer, can deploy from the product. That was the goal all along. A company that never reaches it has not built a product. It has built a staffing agency with an API.

### Building the FDE Team: Reporting Line, Bar, Pairing, and Comp

Category: Leadership
Canonical: https://falkster.com/handbook/fde-org-design

FDEs report to product or engineering, never to sales, because the reporting line sets the scorecard and a bookings scorecard turns an FDE into an SE. The hiring bar is the product engineering bar plus the customer: production code, plus the ability to learn an industry's vocabulary in weeks and earn an operator's trust. It is a senior role; 3 percent of open FDE roles in mid-2026 were early-career, median comp was $185,000, and frontier labs paid above $250,000. Pair every builder with someone who owns adoption, Palantir's Delta and Echo, because the technical build and the organizational change are different skills. Pay variable comp on the customer's outcome and on the renewal it produces, never on new bookings. And size the team to shrink per deployment, because a permanent FDE ratio is a services business.

## The design operating model (5 chapters)

Design moves from producing artifacts to specifying behavior. Every chapter is a consequence of that sentence.

### Design Just Got Promoted

Canonical: https://falkster.com/design/design-just-got-promoted

The claim that AI killed design confuses design with drawing. Drawing got cheap. Deciding did not, and there is more to decide now than at any point in the last twenty years, because every AI feature shipped since 2024 created interface problems that did not exist before. The shift is from producing artifacts to specifying behavior: constraints with failure conditions, a defined range instead of a final state, judgment written into a rubric that scores work without you in the room, and the wrong path owned as a first-class surface. Six operating rules follow, and there is a worksheet at the end you can hand your team on Monday.

### You Cannot Mock a Distribution

Canonical: https://falkster.com/design/you-cannot-mock-a-distribution

When a screen is generated at runtime there is no final state to draw, so the specification is a range instead of a mockup. Three columns: the best case worth aiming at, the worst case you would still ship, and what the product refuses to do. Everything between the second and third is allowed and ships without your review, which is how a designer stops being the bottleneck on a surface that produces a thousand states nobody drew. The worst acceptable column does most of the work. The test that it is real is whether two people sort ten actual outputs the same way.

### Recovery and Trust Repair

Canonical: https://falkster.com/design/recovery-and-trust-repair

Trust repair after an AI is confidently wrong is a designed sequence, and almost nobody has written one. Six steps, in order: detect it, admit it in the same place the wrong answer appeared, contain it by saying what was and was not affected, correct it visibly, name the specific check now in place, and adjust the product's confidence posture on that class of output. Order matters more than eloquence, because good copy in the wrong order still fails. The reason this moment is different from an outage is that the product asserted something and the user acted on it, so they now doubt every earlier answer they did not check. Write the four-sentence script per surface before you need it.

### The Rubric Is the Spec

Canonical: https://falkster.com/design/the-rubric-is-the-spec

A design rubric is the spec, because it is the only form of taste that executes when you are not in the room. Write it by extraction, not invention: grade thirty real outputs on gut feel, find the phrases that repeat in your own reasoning, and keep only the dimensions where failing costs something you can name out loud. Score each dimension pass or fail with a severity attached, because a five-point scale lets two reviewers write 3 for different reasons and believe they agreed. Calibrate with two people on five outputs before you call it real. The template at the end takes one afternoon for the first block, and it is the download this handbook expects to travel furthest.

### Design's Kill List

Canonical: https://falkster.com/design/designs-kill-list

A design ritual belongs on the kill list when it exists to produce a receipt rather than a decision, and the receipt only ever counted because producing it was expensive. Nine fail that test: the pixel-perfect handoff, design review as approval theater, the component library as monument, personas as decoration, the double diamond used as a calendar, fidelity ladders, the design QA pass, the default kickoff workshop, and feedback rounds with no decision rule. Every one of them gets its replacement in the same breath, because killing without replacing is removal, and removal gets reversed inside a quarter. The protocol matters more than the list: name what the ritual protected, ship the replacement first, kill the calendar slot and not just the artifact, and state the condition that would bring it back.

## The agent fleet (44 blueprints)

Each one is a small program that does a slice of product work on its own.
Grouped by the seven stages of the operating system. Ask your assistant to run
one of these directly, or to adapt it to how your team actually works.

### Sense

- **Build Your Competitive Intelligence Agent** · Bi-weekly
  Stop manual competitive research. This agent monitors competitors, tracks customer mentions, and delivers a weekly intel report so you always know your market.
  https://falkster.com/blog/agent-competitive-intel
- **KPI Watchdog Agent: Catch Metric Drops and Ship a Fix Prototype** · Hourly
  A watchdog agent monitors your KPIs hourly, investigates the root cause of any drop, and ships a working prototype fix before you finish your morning coffee.
  https://falkster.com/blog/agent-kpi-watchdog
- **NPS and CSAT Deep Dive Agent** · Daily 8:00 AM + Weekly Monday
  Daily and weekly analysis of your NPS and CSAT scores. Segment breakdowns, feature drivers, and what's actually making your customers happy or frustrated.
  https://falkster.com/blog/agent-nps-csat-analysis
- **Build Your Daily Red Flag Agent** · Daily 9:00 AM
  The first agent every PM should deploy: a morning scan of support, backlog, Slack, and production that tells you exactly what needs your attention today.
  https://falkster.com/blog/agent-red-flag-detection
- **Renewal Risk Agent for Migration Cohorts** · Daily
  Catches the accounts whose operational dashboard says fine but whose qualitative signals say otherwise. Three-dimension scoring, 90 days before renewal.
  https://falkster.com/blog/agent-renewal-risk-during-migration
- **Automate Support Pattern Detection** · Daily 8:00 AM
  A daily agent that extracts signals from your support queue: which problems are rising, which segments are hurting, and what real trend lies beneath the noise.
  https://falkster.com/blog/agent-support-signal-processing

### Discover

- **Customer Segmentation Agent** · Weekly Monday
  Weekly updates to your customer segments and personas. Cohort analysis, segment evolution, and behavioral groupings - automatically maintained.
  https://falkster.com/blog/agent-customer-segmentation
- **Customer Interview Synthesis Agent** · Weekly Wednesday
  Weekly automated synthesis of your customer interviews. Key themes, hypotheses, and actionable insights - without the tedious manual work.
  https://falkster.com/blog/agent-interview-synthesis
- **Automated Customer Journey Mapping** · Bi-weekly Thursday
  Build rich customer journey maps bi-weekly from session replays, support patterns, and research data. See exactly where users get stuck.
  https://falkster.com/blog/agent-journey-mapping
- **Build Your Market Intelligence Agent** · Bi-weekly
  Bi-weekly deep dive into your competitive landscape. Win/loss analysis, market shifts, feature parity, and strategic recommendations so you're never surprised.
  https://falkster.com/blog/agent-market-intel
- **Build Your Weekly Ops Digest Agent** · Weekly Monday 8:00 AM
  Daily agents catch fires. This weekly digest spots the trends: recurring issues, degrading resolution times, and systemic problems that daily reports miss.
  https://falkster.com/blog/agent-weekly-ops-digest

### Decide

- **Testable Assumptions Tracker Agent** · Weekly Friday
  Convert opportunities into testable assumptions. Track validation status weekly. Know which assumptions are holding up your roadmap.
  https://falkster.com/blog/agent-assumption-tracker
- **Build Your Customer Commitment Agent** · Bi-weekly
  Sales promised a feature by Q2. CS said 'next month.' This agent tracks every commitment, cross-references the roadmap, and flags what is overdue.
  https://falkster.com/blog/agent-customer-commitments
- **Build Your Engineering Capacity Agent** · Weekly Monday 8:00 AM
  Engineering capacity agent: flags overloaded engineers, catches PTO gaps weeks early, and tells you if the roadmap won't fit before the sprint starts.
  https://falkster.com/blog/agent-engineering-capacity
- **OKR Progress and Prediction Agent** · Daily 4:00 PM + Weekly Friday
  Daily OKR tracking with outcome scoring. Weekly confidence predictions for OKR achievement. See what's on track and what needs intervention.
  https://falkster.com/blog/agent-okr-tracker
- **Opportunity Prioritization and Synthesis Agent** · Weekly Friday
  Weekly synthesis of all your DISCOVER agent outputs into a prioritized opportunity stack. OST-ready opportunities with impact estimates and dependencies.
  https://falkster.com/blog/agent-opportunity-prioritization
- **Pricing Migration Tracker Agent** · Daily
  Watches every migration account daily, classifies drift type, and recommends a specific intervention. Catch quiet drifters in week one, not week six.
  https://falkster.com/blog/agent-pricing-migration-tracker
- **Build Your Roadmap Progress Agent** · Daily 9:00 AM
  Your roadmap says in progress but engineering hasn't touched it in weeks. This agent cross-references your backlog against actual dev activity daily.
  https://falkster.com/blog/agent-roadmap-tracker
- **Automated Sprint Planning Agent** · Weekly Monday
  Convert prioritized opportunities into user stories and sprint plans every Monday. Accounts for team capacity, tech debt, and dependencies.
  https://falkster.com/blog/agent-sprint-planning

### Build

- **Auto Bugfix Agent: Zendesk Ticket to Reviewable PR in 11 Minutes** · On-demand + Daily 8:00 AM
  An AI agent that reproduces customer-reported bugs, locates the broken code, writes the fix, adds a regression test, and opens a reviewable PR.
  https://falkster.com/blog/agent-auto-bugfix
- **Build Your Documentation Gap Agent** · Daily 7:00 AM
  Features ship without docs and customers find the gaps. This agent scans tickets and the release calendar to catch documentation gaps before they become fires.
  https://falkster.com/blog/agent-documentation-gaps
- **Instant Prototype Agent: Customer Request to Prototype in Minutes** · On-demand + Daily 9:00 AM
  An AI agent that turns customer feature requests into working prototypes. Deploys a preview URL, opens a PR, files a Linear ticket, drafts a Notion doc.
  https://falkster.com/blog/agent-instant-prototype
- **Automated PRD Generator** · Weekly Monday
  Convert prioritized opportunities into PRDs automatically. Drafts based on research context, design specs, and technical requirements.
  https://falkster.com/blog/agent-prd-generator
- **Build Your Product Ops Agent** · Daily 9:00 AM
  Daily product ops report ranked by ARR at risk: top complaints, feature gaps, usage anomalies, and at-risk customers synthesized from five data streams.
  https://falkster.com/blog/agent-product-ops
- **Build Your Team Triage Agent** · Daily 9:00 AM and 4:00 PM
  Your #team-product channel is a firehose. This agent reads every message, categorizes issues, assigns owners, and surfaces what's unanswered, twice daily.
  https://falkster.com/blog/agent-team-triage
- **Tech Debt Impact and Prioritization Agent** · Weekly Monday
  Weekly analysis of tech debt impact on velocity. Which debt items actually matter? What's the priority order to reclaim speed?
  https://falkster.com/blog/agent-tech-debt-analyzer

### Ship

- **Build Your GTM Release Monitoring Agent** · Daily 7:00 AM
  Track if features are ready to sell, support, and succeed. Monitors discovery compliance, beta standards, and GTM materials daily across five readiness gates.
  https://falkster.com/blog/agent-gtm-monitoring
- **Launch Comms Agent: Six Channels Generated on Every Release** · On production deploy + Thursday 10:00 sweep
  AI agent that generates website copy, LinkedIn post, customer email, in-app banner, changelog, and X thread when a feature ships. Six channels, four minutes.
  https://falkster.com/blog/agent-launch-comms
- **Build Your Release Checker Agent** · Weekly Thursday
  The final gate before you ship. This Thursday agent verifies QA, docs, GTM materials, and sign-offs, giving a clear go/no-go for each feature.
  https://falkster.com/blog/agent-release-checker
- **Automated Release Documentation Agent** · Weekly Wednesday + per-release
  Auto-generate release notes, customer comms, and doc updates as features ship. Saves hours on documentation and keeps comms consistent.
  https://falkster.com/blog/agent-release-documentation
- **Build Your Release Readiness Agent** · Weekly Wednesday
  Prevents launch disasters. Checks PRDs, release notes, feature flags, and GTM readiness every Wednesday so nothing ships without proper preparation.
  https://falkster.com/blog/agent-release-readiness

### Measure

- **Direction Dashboard Agent** · Daily
  Compiles seven leading indicators every morning, predicts outcomes 4-8 weeks ahead, runs a weekly Goodhart audit. Direction at AI-native velocity.
  https://falkster.com/blog/agent-direction-dashboard
- **Feature Adoption Tracking Agent** · Daily 4:00 PM + Weekly Monday
  Daily adoption curves for new features. Identify stuck cohorts and recommend interventions before adoption stalls.
  https://falkster.com/blog/agent-feature-adoption
- **Margin Watch Agent** · Daily (operational) / Weekly (CPO digest)
  Tracks gross margin per outcome daily, classifies compression cause, and forecasts the Jevons cliff. Pricing reviews stop being post-mortems.
  https://falkster.com/blog/agent-margin-watch
- **Build Your PM Issues Agent** · Daily 7:00 AM
  Catch slipping features, overdue customer promises, and cross-team dependency failures before they become fires. Scans roadmap and commitments every morning.
  https://falkster.com/blog/agent-pm-issues
- **Build Your Product Health Dashboard Agent** · Weekly Tuesday
  Go deeper than daily metrics. Weekly agent analyzing feature adoption curves, cohort retention, engagement depth, and performance trends to shape strategy.
  https://falkster.com/blog/agent-product-dashboard
- **Build Your Product Health Agent** · Daily 4:00 PM
  Daily 4 PM pulse check: engagement, activation, performance, and support trends synthesized into one story. Know exactly where your product stands each day.
  https://falkster.com/blog/agent-product-health
- **Signal-to-Ship Cycle Time Agent: Measure PM Velocity Across 7 Stages** · Daily 9:00 AM + Weekly Monday
  An AI agent that tracks PM cycle time across the 7 stages of the Product Operating System (Sense to Amplify) and surfaces the weekly bottleneck.
  https://falkster.com/blog/agent-signal-to-ship
- **Win/Loss Analysis Agent** · Bi-weekly Tuesday
  Bi-weekly analysis of won and lost deals. Extract product insights from sales outcomes. What's winning business, and what's holding you back?
  https://falkster.com/blog/agent-win-loss-analysis

### Amplify

- **Board Narrative Drafter Agent** · Weekly (drafts) / Quarterly (final)
  The most expensive document in product management compiled while you sleep. Quarterly board updates from data sources, comp set, and last quarter's narrative.
  https://falkster.com/blog/agent-board-narrative-drafter
- **Build Your Daily Focus Agent** · Daily 7:00 AM
  Your AI chief of staff: reads your calendar, scans Slack, checks the roadmap, and delivers the 3 things that actually matter today before your first meeting.
  https://falkster.com/blog/agent-daily-focus
- **Build Your Executive Report Agent** · Weekly Monday 1:00 PM
  Auto-generated Monday leadership brief: roadmap status, top wins, key risks, and metrics in a 3-minute read. Your CEO walks in already prepared.
  https://falkster.com/blog/agent-executive-report
- **Retrospective Synthesis and Learning Agent** · Weekly Friday
  Extract learnings from sprint retros automatically. Update playbooks, surface patterns, and drive continuous improvement.
  https://falkster.com/blog/agent-retrospective-synthesis
- **Stakeholder Communication Agent** · Weekly Friday + Monthly
  Generate tailored updates for different audiences: exec summaries, board updates, investor reports, team briefings. All from the same data.
  https://falkster.com/blog/agent-stakeholder-communication

## The toolkit (114 templates)

Working documents, not illustrations. Ask your assistant to fill one in with
your own situation rather than to describe what it contains.

- **AARRR Dashboard: Metric Definitions, Thresholds, and Investigation Playbooks** (3977 words) · https://falkster.com/toolkit/aarrr-dashboard
- **Board Narrative Drafter Agent** (1040 words) · https://falkster.com/toolkit/agent-board-narrative-drafter
- **Weekly Competitive Intelligence Agent** (2470 words) · https://falkster.com/toolkit/agent-competitive-intel
- **Customer Commitment Agent, Complete Prompt** (2586 words) · https://falkster.com/toolkit/agent-customer-commitments
- **Daily Focus Agent, Complete Prompt** (1867 words) · https://falkster.com/toolkit/agent-daily-focus
- **The Agent Department Org-Design Worksheet** (645 words) · https://falkster.com/toolkit/agent-department-org-design-worksheet
- **Direction Dashboard Agent** (1128 words) · https://falkster.com/toolkit/agent-direction-dashboard
- **Documentation Gap Detection Agent** (2210 words) · https://falkster.com/toolkit/agent-documentation-gaps
- **Engineering Capacity Agent, Complete Prompt** (2821 words) · https://falkster.com/toolkit/agent-engineering-capacity
- **Weekly Executive Report Agent** (1403 words) · https://falkster.com/toolkit/agent-executive-report
- **GTM Release Monitoring Agent** (2686 words) · https://falkster.com/toolkit/agent-gtm-monitoring
- **KPI Watchdog Agent** (1346 words) · https://falkster.com/toolkit/agent-kpi-watchdog
- **Margin Watch Agent** (1198 words) · https://falkster.com/toolkit/agent-margin-watch
- **Market Intelligence Agent** (2252 words) · https://falkster.com/toolkit/agent-market-intel
- **Daily PM Issues Report Agent** (1942 words) · https://falkster.com/toolkit/agent-pm-issues
- **Pricing Migration Tracker Agent** (1215 words) · https://falkster.com/toolkit/agent-pricing-migration-tracker
- **Product Health Dashboard Agent** (1839 words) · https://falkster.com/toolkit/agent-product-dashboard
- **Product Health Agent, Complete Prompt** (2540 words) · https://falkster.com/toolkit/agent-product-health
- **Daily Product Ops Agent** (2791 words) · https://falkster.com/toolkit/agent-product-ops
- **Daily Red Flag Detection Agent** (1839 words) · https://falkster.com/toolkit/agent-red-flag-detection
- **Release Checker Agent** (2354 words) · https://falkster.com/toolkit/agent-release-checker
- **Release Readiness Agent** (2879 words) · https://falkster.com/toolkit/agent-release-readiness
- **Renewal Risk Agent for Migration Cohorts** (1157 words) · https://falkster.com/toolkit/agent-renewal-risk-during-migration
- **Roadmap Progress Tracker Agent** (2531 words) · https://falkster.com/toolkit/agent-roadmap-tracker
- **Signal-to-Ship Cycle Time Agent** (2103 words) · https://falkster.com/toolkit/agent-signal-to-ship
- **Team Triage Agent, Complete Prompt** (2187 words) · https://falkster.com/toolkit/agent-team-triage
- **Weekly Ops Digest Agent** (1514 words) · https://falkster.com/toolkit/agent-weekly-ops-digest
- **AI Eval Starter Kit** (1469 words) · https://falkster.com/toolkit/ai-eval-starter-kit
- **AI Noise vs Signal Audit** (1665 words) · https://falkster.com/toolkit/ai-noise-vs-signal-audit
- **The Board Narrative Slide Outline** (655 words) · https://falkster.com/toolkit/board-narrative-slide-outline
- **Board Deck Product Section: 7-Slide Skeleton** (1285 words) · https://falkster.com/toolkit/board-product-section
- **Build Agent Stack — Starter Pack** (836 words) · https://falkster.com/toolkit/build-agent-stack-starter
- **20-Minute Customer Call Triage Agent: Full Recipe** (1129 words) · https://falkster.com/toolkit/call-triage-agent
- **The Cannibalization Decision Tree** (949 words) · https://falkster.com/toolkit/cannibalization-decision-tree
- **Competitive Intelligence Prompts** (718 words) · https://falkster.com/toolkit/competitive-intel-prompts
- **CPO Coalition Map Worksheet (First 90 Days)** (1050 words) · https://falkster.com/toolkit/cpo-coalition-map-worksheet
- **The CPO Coalition Map** (998 words) · https://falkster.com/toolkit/cpo-coalition-map
- **CPO Day-90 Readout Skeleton** (1236 words) · https://falkster.com/toolkit/cpo-day-90-readout
- **CPO Listening Tour Interview Guide** (1235 words) · https://falkster.com/toolkit/cpo-listening-tour-guide
- **CPO Trust Ledger + Claim Ledger** (910 words) · https://falkster.com/toolkit/cpo-trust-ledger
- **Decision Log** (1301 words) · https://falkster.com/toolkit/decision-log
- **The Constraint Template** (1303 words) · https://falkster.com/toolkit/design-constraint-template
- **The Design Decision Record** (900 words) · https://falkster.com/toolkit/design-decision-record
- **The Disclosure Ladder** (875 words) · https://falkster.com/toolkit/design-disclosure-ladder
- **The Evidence Register** (895 words) · https://falkster.com/toolkit/design-evidence-register
- **The Exposure Grid** (938 words) · https://falkster.com/toolkit/design-exposure-grid
- **The Failure Inventory Worksheet** (1033 words) · https://falkster.com/toolkit/design-failure-inventory
- **Your First Ninety Days** (1063 words) · https://falkster.com/toolkit/design-first-90-days
- **The Interview Loop for Judgment** (1054 words) · https://falkster.com/toolkit/design-interview-loop
- **The Twelve-Week Judgment Curriculum** (990 words) · https://falkster.com/toolkit/design-judgment-curriculum
- **Design's Kill List** (1108 words) · https://falkster.com/toolkit/design-kill-list
- **Latency States by Duration Band** (923 words) · https://falkster.com/toolkit/design-latency-states
- **The Measurement Starter Set** (940 words) · https://falkster.com/toolkit/design-measurement-starter-set
- **The One-Slide Argument** (992 words) · https://falkster.com/toolkit/design-one-slide-argument
- **The Operating Cadence, One Page** (902 words) · https://falkster.com/toolkit/design-operating-cadence
- **The Design Operating Rules** (1078 words) · https://falkster.com/toolkit/design-operating-rules
- **Three Design Org Patterns by Stage** (985 words) · https://falkster.com/toolkit/design-org-patterns
- **The Ownership Map** (846 words) · https://falkster.com/toolkit/design-ownership-map
- **The New Portfolio Structure** (898 words) · https://falkster.com/toolkit/design-portfolio-structure
- **The Range Spec** (1016 words) · https://falkster.com/toolkit/design-range-spec
- **Refusal Copy Patterns** (924 words) · https://falkster.com/toolkit/design-refusal-patterns
- **The Review Protocol, Thirty Minutes** (1019 words) · https://falkster.com/toolkit/design-review-protocol
- **The Design Rubric Template** (1140 words) · https://falkster.com/toolkit/design-rubric-template
- **The Trust Repair Sequence** (1054 words) · https://falkster.com/toolkit/design-trust-repair
- **The Uncertainty Pattern Library** (1085 words) · https://falkster.com/toolkit/design-uncertainty-patterns
- **The Undo Checklist** (992 words) · https://falkster.com/toolkit/design-undo-checklist
- **The Weekly Research Cadence** (1090 words) · https://falkster.com/toolkit/design-weekly-research-cadence
- **Discovery Agent Stack — Starter Pack** (776 words) · https://falkster.com/toolkit/discovery-agent-stack-starter
- **The Dual Transformation Operating Cadence** (492 words) · https://falkster.com/toolkit/dual-transformation-operating-cadence
- **The Eval Rubric Template** (728 words) · https://falkster.com/toolkit/eval-rubric-template
- **The Eval Template** (621 words) · https://falkster.com/toolkit/eval-template
- **Exec Update** (1016 words) · https://falkster.com/toolkit/exec-update
- **Falkster Design System, for Claude** (2609 words) · https://falkster.com/toolkit/falkster-design-system
- **Feedback Pipeline Setup Guide** (517 words) · https://falkster.com/toolkit/feedback-pipeline-setup
- **The Guerrilla PM Playbook: Operating Without a Team** (2549 words) · https://falkster.com/toolkit/guerrilla-pm-playbook
- **Principal Product Builder** (866 words) · https://falkster.com/toolkit/jd-principal-product-builder
- **The Product Builder Job Ladder** (1543 words) · https://falkster.com/toolkit/jd-product-builder-ladder
- **Product Builder** (738 words) · https://falkster.com/toolkit/jd-product-builder
- **Senior Product Builder** (743 words) · https://falkster.com/toolkit/jd-senior-product-builder
- **Staff Product Builder** (798 words) · https://falkster.com/toolkit/jd-staff-product-builder
- **The Margin Recovery Curve Model** (699 words) · https://falkster.com/toolkit/margin-recovery-curve-model
- **Measurement Agent Stack — Starter Pack** (935 words) · https://falkster.com/toolkit/measurement-agent-stack-starter
- **OKR Writing Guide: Templates, Examples, and Anti-Pattern Checklist** (2319 words) · https://falkster.com/toolkit/okr-writing-guide
- **Old PM vs Product Builder — The Ledger** (745 words) · https://falkster.com/toolkit/old-pm-vs-product-builder
- **[Product or Feature name] — One-Page Brief** (364 words) · https://falkster.com/toolkit/one-page-brief
- **One-Pager Template** (550 words) · https://falkster.com/toolkit/one-pager-template
- **PM Agent Stack — The Master Install Checklist** (697 words) · https://falkster.com/toolkit/pm-agent-stack-install-checklist
- **PM AI Prompts Reference Card** (520 words) · https://falkster.com/toolkit/pm-ai-prompts
- **The Customer Signal Monitoring Playbook** (1537 words) · https://falkster.com/toolkit/pm-customer-signal-playbook
- **PM Bet One-Pager + Day-90 Note** (951 words) · https://falkster.com/toolkit/pm-day-90-note
- **PM Decision Archaeology Worksheet** (1018 words) · https://falkster.com/toolkit/pm-decision-archaeology
- **PM Manager Contract (One Page)** (856 words) · https://falkster.com/toolkit/pm-manager-contract
- **Make the Case: PMs Shipping Production Code, Pilot Proposal (Regulated Environment Variant)** (1709 words) · https://falkster.com/toolkit/pm-pitch-doc-enterprise
- **Make the Case: PMs Shipping Production Code, Pilot Proposal** (1327 words) · https://falkster.com/toolkit/pm-pitch-doc-standard
- **PLANNING: [Feature Name]** (595 words) · https://falkster.com/toolkit/pm-planning-template
- **PM PR Review Skill** (1327 words) · https://falkster.com/toolkit/pm-pr-review-skill
- **PM Signal Map Worksheet** (853 words) · https://falkster.com/toolkit/pm-signal-map
- **Claude Skills Templates for PMs** (969 words) · https://falkster.com/toolkit/pm-skills-templates
- **The PM-to-CPO 12-Month Roadmap** (858 words) · https://falkster.com/toolkit/pm-to-cpo-12-month-roadmap
- **PR/FAQ Template** (1132 words) · https://falkster.com/toolkit/pr-faq
- **The Pricing Migration Quarterly Tracker** (632 words) · https://falkster.com/toolkit/pricing-migration-quarterly-tracker
- **Product Trio Setup Guide** (1965 words) · https://falkster.com/toolkit/product-trio-guide
- **The Prototype Brief** (392 words) · https://falkster.com/toolkit/prototype-brief
- **60-Minute Prototype Workflow** (546 words) · https://falkster.com/toolkit/prototype-workflow
- **The Punctuated Discovery Sprint Template** (660 words) · https://falkster.com/toolkit/punctuated-discovery-sprint-template
- **QBR Deck Outline** (1434 words) · https://falkster.com/toolkit/qbr-deck-outline
- **Quarterly Business Review Template: Complete Structure, Talking Points, and Metric Framework** (3688 words) · https://falkster.com/toolkit/quarterly-business-review
- **Sprint Retro Toolkit** (868 words) · https://falkster.com/toolkit/retro-checklist
- **Complete Data Source Setup Guide for Claude** (3965 words) · https://falkster.com/toolkit/setup-guide-claude
- **Complete Data Source Setup Guide for OpenClaw** (4546 words) · https://falkster.com/toolkit/setup-guide-openclaw
- **Stakeholder Update Templates** (485 words) · https://falkster.com/toolkit/stakeholder-update-agent
- **Product Strategy Memo Template (Pyramid Principle)** (1202 words) · https://falkster.com/toolkit/strategy-memo
- **The Sunset Communication Template** (955 words) · https://falkster.com/toolkit/sunset-communication-template
- **Weekly Review Checklist** (371 words) · https://falkster.com/toolkit/weekly-review-checklist

## The answer library (198 questions)

Narrow questions with direct answers. These are the positions, stated plainly.

### What is Teresa Torres's three-step approach to AI evals?

Three steps, published on Product Talk on September 2, 2026. One, error analysis: read the model's outputs yourself, at volume, and write down the two or three mistakes that recur, ranked by what they cost a customer. Two, match each error to an eval type: a code assertion when the check is deterministic, a golden dataset when there is one right answer, an LLM-as-a-judge when the judgment is semantic, and customer feedback signals such as regeneration and editing for real usage, always trying the cheapest first and using it to filter before the expensive one. Three, the experimentation loop: fixed production-like inputs, a baseline, one change, compare across every eval, iterate. Her rule underneath all of it is that correctness is the product team's definition and never the vendor's. The step I add is the gate: a passing eval on a reversible action buys autonomy, a passing eval on an irreversible one buys a draft and a human, and model confidence never decides.

https://falkster.com/answers/what-is-teresa-torres-three-step-approach-to-ai-evals

### Do forward deployed engineers scale, or is it just consulting?

Only if two things happen, and most companies adopting the title in 2026 have set up neither. What the FDE builds for one customer has to be generalized so later customers need less of an engineer, and the pricing has to capture the outcome the engineer guaranteed. Palantir ran the model for a decade with more FDEs than product engineers until 2016 and, per Plank's figure, roughly $1.5 million in revenue per employee with almost no traditional sales force, because the FDEs were the sales motion and what they built became Foundry. Sacra flagged implementation intensity as Sierra's biggest scaling risk, and Sierra's answer was to move workflow building into a no-code studio so each deployment needs less engineering. The buyer's objection is real: Andrew Ng warns that a vendor's FDEs reduce optionality, and Gartner expects enterprises to abandon FDE-heavy engagements on cost. The fix is contractual: the observation record and every configuration belong to the customer and port with them. The signal to watch is FDE hours per deployment for the same customer class. Falling means the loop is closing. Flat means you are running a services firm and calling it product.

https://falkster.com/answers/do-forward-deployed-engineers-scale

### Does zero data retention mean the vendor keeps nothing from your AI agents?

No. Zero data retention is a promise about the model provider: your business data is sent to answer a question and the model company keeps none of it. Marc Benioff made exactly that promise at Dreamforce 26, and Claudeforce runs on it. What it does not cover is the observation record generated wherever the agents run: which actions were taken, which a human reversed and how fast, what was escalated, which rule fired, and how the outcome closed. That record is the only asset in an agent deployment that compounds, and it accrues to whoever operates the agents, which in the Salesforce case is Salesforce. Patrick Stokes said as much on stage when he said the company's value is the trust customers put in it to hold their data. The question to ask a vendor is not whether the model keeps your data. It is who owns the trace of what the agents did in your org, who can read it, and whether it leaves with you.

https://falkster.com/answers/does-zero-data-retention-mean-the-vendor-keeps-nothing

### How is a forward deployed engineer different from a sales engineer?

When their accountability ends. A sales engineer supports the sale with demos, proofs of concept, and technical answers, and is done when the customer signs. A forward deployed engineer's job begins at the signature: they build inside the customer's environment, integrate with real data and workflows, and are measured on whether the outcome the customer bought actually shows up after go-live. Every other difference, on-site time, production code, staying through the first months, follows from that one line. The same test sorts the neighbors: a solutions architect is accountable for the design, a consultant for the recommendation, professional services for the fixed spec, developer relations for reach. Only the FDE is on the hook for the result. Hire an SE when you needed an FDE and you get beautiful pilots and no production, which is my read on most of the 95 percent of enterprise AI pilots MIT NANDA found with no measurable impact.

https://falkster.com/answers/how-is-a-forward-deployed-engineer-different-from-a-sales-engineer

### What is a forward deployed engineer?

A forward deployed engineer is an engineer employed by a software vendor, embedded inside a customer's organization, and accountable for the outcome the software produces there. Palantir popularized the role around 2008 after learning that software shipped without its engineer did not get adopted; its embedded engineers were called Deltas, each paired with an Echo who owned adoption and politics, and until 2016 the company employed more FDEs than product engineers. In 2026 the role is everywhere: OpenAI raised more than $4 billion for a deployment company in May, Anthropic launched a $1.5 billion enterprise joint venture the same month, and Google Cloud, Databricks, Salesforce, and Sierra all run a version. The one thing that survived every vintage of the role, per Natalie Meurer of Sierra, is that every FDE is accountable to the customer. That accountability is what separates it from a sales engineer, whose job ends at the signature. My read is that it is the Product Builder role arriving from the services side: same six skills, different badge.

https://falkster.com/answers/what-is-a-forward-deployed-engineer

### When can you actually enforce a rule on an AI agent?

When the decision is observable at a chokepoint every call already passes through, can be decided without the agent's reasoning, and is cheap to reverse if wrong. Two of the three and you can build a gate. Spotify's shunt plugin is the clean case: a hook blocks any file read over 350 lines because a path and a line count are visible before the read happens, but the boilerplate-writing path stays advisory because whether a piece of code is boilerplate or design is a judgment the gate cannot make without doing the reasoning it was meant to avoid. Their first attempt, rules in CLAUDE.md, was ignored. A rule that fails the test is a preference, however firmly it is written, and the honest move is to know which one you are holding.

https://falkster.com/answers/when-can-you-enforce-a-rule-on-an-ai-agent

### Is llms.txt enough to make a site useful to AI assistants?

No, because llms.txt is an index and not content. It tells a crawler what exists and where, which helps discovery and does nothing for the person who wants your actual thinking inside their own assistant. Anthropic, Cloudflare, Supabase, and Vercel all publish one, and the public directories tracking adoption list over 780 sites doing the same. Publishing the index is now table stakes. The part almost nobody does is publishing the substance: a downloadable corpus of your claims and summaries that someone loads into a project once and works from offline.

https://falkster.com/answers/is-llms-txt-enough-for-ai-search

### What should I put in a Claude Project for product management?

Put in the things a general model does not have: positions it can defend, artifacts it can fill in, and links it can fetch when a decision turns on the detail. Skip definitions and frameworks explained in the abstract, because the model already knows those better than your notes do, and adding them just spends context to reproduce the average of the internet. The practical test is whether a line in your knowledge base would change an answer. A page defining the RICE framework will not. A rule saying you kill anything that has not shipped a measurable outcome in two quarters will.

https://falkster.com/answers/what-should-i-put-in-a-claude-project-for-product-management

### What is Marty Cagan's fresh definition of the product role?

Strictly, he does not give a new one, and that is his argument. In the August 2026 SVPG piece Cagan takes three skills he credits to Benedict Evans, recognizing the general problem behind specific pain, being able to create a solution rather than merely use one, and understanding what a solution means across the whole company, and maps them onto the framework SVPG has used for two decades: problem discovery, value risk, and viability risk. His conclusion is that the definition survived AI intact, because AI changes the thresholds and not the problem. The skills do survive. What does not survive is the assumption that they are one job in roughly equal parts, because the three have completely different exposure to what just happened.

https://falkster.com/answers/what-is-marty-cagans-definition-of-the-product-role

### Can you retrofit context for AI agents?

No. Every other layer in an agent stack can be added after the fact, because every other layer is a read of something that already exists. A control plane reads systems you already own, model routing is a config change, governance can be wrapped around agents that are already running. Context is different: the reason a decision was made exists only at the moment someone makes it, and only that person can record it. Logs give you what happened, schemas give you how, and neither reconstructs why. If it wasn't written down then, it is not recoverable, which makes where your team records reasoning a decision you are making this quarter rather than at renewal.

https://falkster.com/answers/can-you-retrofit-context-for-ai-agents

### Why did Miro sell for only 2.3x ARR?

Because a multiple prices the next five years, and Miro's core artifact stopped being readable by the thing that now does the next step. Bending Spoons is acquiring Miro at a $1.355B enterprise value on roughly $600M ARR, about 2.3x, against a $17.5B mark from January 2022. The company is profitable, holds roughly $435M in net cash, and takes close to 90% of revenue from business and enterprise customers. Fundamentals like that clear well above 2.3x routinely, so compression explains the range but not the position inside it. A whiteboard imposes no schema, which is exactly why nothing downstream can consume its output, and every decision made on a board gets re-entered by hand into a tool that has a data model.

https://falkster.com/answers/why-did-miro-sell-for-2-3x-arr

### Why is a polished mockup no longer evidence of conviction?

Because the cost of looking finished collapsed and the cost of being right did not move. For twenty years reviewers read fidelity as a proxy for conviction: grayscale meant early thinking so you attacked the concept, pixel-perfect meant someone had lived with the problem for two weeks so you moved to states and edge cases. A polished screen with real copy and a working interaction can now be forty minutes old. The proxy still feels true, which is why reviews now fail in one predictable direction: the work that looks most done gets the shallowest critique. The fix is to move the signal off the artifact and into the room, by having the presenter state a confidence level out loud before the screen share.

https://falkster.com/answers/why-is-a-polished-mockup-not-evidence-of-conviction

### Should you be the platform underneath or the application on top?

Decide it by asking what has to happen for you to get paid again next year, not by which layer sounds more defensible. If renewal depends on a human choosing to come back, you are a surface, and your job is habit. If it depends on a system calling you, you are a substrate, and your job is reliability, permission fidelity, and coverage. Glean and Alation are currently doing both, which is optionality bought with a balance sheet most companies do not have. Plumbing has no wDAU/wMAU ratio, because nobody opens plumbing.

https://falkster.com/answers/should-you-be-a-platform-or-an-application

### What did Salesforce's Q2 FY27 earnings say about agentic AI?

That the model layer is capturing value faster than the applications on top of it, and that the pricing unit is moving from seats to work. Adjusted EPS of $5.90 matched consensus; the 119% GAAP jump came from a $2.6 billion strategic investment gain tied to Salesforce's Anthropic stake, now roughly two thirds of its investment portfolio. Salesforce also announced Claudeforce, shipping its CRM as a plugin inside Claude. But Agentforce ARR of $1.5 billion is only about 3% of a $46 billion guide and the seats did not collapse, so this is an attach motion, not a replacement cycle.

https://falkster.com/answers/what-did-salesforce-q2-fy27-say-about-agentic-ai

### What are bottom-up evals, and why can't an LLM write them?

Bottom-up evals are the checks you discover by reading a pile of real outputs and noticing what bothers you. Top-down evals are the ones you can write from the task description alone, which is why a model drafts them well: a top-down eval is the brief restated in a new format, and translation is what these models do best. An LLM cannot produce bottom-up evals because they require having read the outputs and having a stake in whether they are good. The conversion rule is one line: every gut reaction becomes a yes/no a grader can answer without judgment, or it gets dropped.

https://falkster.com/answers/what-are-bottom-up-evals

### Why do customers buy your product but not renew it?

Because the reason a buyer approved you and the reason anyone keeps using you have come apart, and most teams measure only the first. Cost, risk, or a compliance deadline is the permission that gets the purchase order signed. Return behavior is what makes the account renew and expand. Glean's $300M ARR run was sold on cutting AI spend, but its 45% wDAU/wMAU ratio, double the enterprise SaaS norm, is what kept expansion compounding across departments. Instrument both layers separately, and treat unprompted return rate as the renewal forecast.

https://falkster.com/answers/why-do-customers-buy-but-not-renew

### How do you design a screen that is generated at runtime?

You specify the range instead of the screen. Three columns: the best case worth aiming at, the worst case you would still ship, and what the product refuses to do. Everything between the second and third columns is allowed and does not need your review, which is how a designer stops being the bottleneck on a surface that produces a thousand states nobody drew. A mockup of a generated surface is one sample from a distribution presented as if it were the decision.

https://falkster.com/answers/how-do-you-design-a-screen-that-is-generated-at-runtime

### How do you write a design rubric that scores work without you?

You extract it, you do not invent it. Grade thirty real outputs from the surface on gut feel with one sentence of reasoning each and no criteria written yet, then take your candidate dimensions from the phrases that repeat. Keep a dimension only if you can name what failing it costs in money, time, trust, harm, or rework, which usually leaves three or four instead of eight. Write each one as a binary pass and fail test with a severity attached, then calibrate: two people score the same five outputs independently until they land within one disagreement.

https://falkster.com/answers/how-do-you-write-a-design-rubric

### Which design rituals should you kill, and what replaces them?

Kill the rituals that produce a receipt rather than a decision, where the receipt only ever counted because producing it was expensive. That test catches nine: pixel-perfect handoff, approval-theater review, the component library as catalogue, laminated personas, the double diamond used as a calendar, mandatory fidelity ladders, the design QA pass, the default kickoff workshop, and feedback rounds with no decision rule. Every one needs its replacement named and running before the kill, because killing without replacing is removal and removal gets reversed inside a quarter.

https://falkster.com/answers/what-design-rituals-should-you-kill

### What do you do in the thirty seconds after your AI is confidently wrong?

Run a sequence, in order: detect it, admit it in the same place the wrong answer appeared, contain it by saying what was and was not affected, correct it visibly, name the specific check now in place, then adjust the product's confidence posture on that class of output. Order matters more than eloquence, because good copy delivered out of order still fails. Write the four-sentence script per surface before you need it, since nobody writes good copy during an incident.

https://falkster.com/answers/what-do-you-do-after-your-ai-is-confidently-wrong

### What do designers do when AI generates the screens?

They specify behavior instead of producing artifacts. When a team can put twelve credible directions on the wall by Wednesday, the scarce work is saying why this one and not those eleven, in language a team and an agent can both execute. That means writing constraints with failure conditions attached, defining the range instead of the final state, moving taste into a rubric that scores work without you in the room, and owning the wrong path: recovery, disclosure, confidence, correction. That work used to be roughly fifteen percent of a designer's week, hidden inside the making. It is the whole week now.

https://falkster.com/answers/what-do-designers-do-when-ai-generates-the-screens

### Why did SpaceX buy Cursor for $60 billion?

SpaceX agreed in June 2026 to acquire Cursor's parent, Anysphere, for $60 billion in stock, and closed on August 14, days after its own record IPO and its merger with xAI. Two reasons, and the second is the one nobody says out loud. First, Cursor was the crown jewel of AI coding: the fastest company ever to $100M ARR (January 2025), past $1 billion annualized by November 2025, with more than half the Fortune 500 as customers, all on developer word of mouth. Second, the deal fixes Cursor's one structural weakness. Cursor ran on outside frontier models, mainly Anthropic's Claude, so a supplier could become a competitor overnight. Pairing it with xAI gives Cursor a captive model supply it can never be cut off from.

https://falkster.com/answers/why-did-spacex-buy-cursor

### Why is Harvey AI worth $11 billion?

Harvey raised $200 million at an $11 billion valuation in March 2026, co-led by GIC and Sequoia. It reached $100M ARR about three years after founding, $195M by the end of 2025, and roughly $350M by July 2026, selling to the hardest buyer there is: elite law firms, with more than 142,000 lawyers and half the Am Law 100 as customers. The valuation is not a bet on the model. It is a bet on vertical depth. Harvey owns the legal-specific workflows, the trust that lets a partner stake their name on the output, and the distribution into Big Law, none of which a frontier lab can ship in a weekend. Founded in 2022 by ex-litigator Winston Weinberg and ex-DeepMind researcher Gabriel Pereyra, it built for law specifically rather than horizontally.

https://falkster.com/answers/why-is-harvey-ai-worth-11-billion

### Does the IKEA vs Salesforce story prove AI automation fails?

No, and it argues the opposite. Salesforce's Agentforce hit its target (about half of conversations handled, cases down, support cost down 17%), and IKEA only freed capacity to reskill 8,500 people because Billie deflected 47% of inquiries. The comparison is junk: the two used different technology a decade apart (IKEA's Billie was a 2021 intent router, Salesforce's Agentforce a 2025 agent system), the 47% is a containment rate misread as a quality score, and the rehiring-at-1.5x and Benioff-regret claims are unsourced. Both companies actually ran the same play: move deflected-work staff toward revenue.

https://falkster.com/answers/does-the-ikea-story-prove-ai-fails

### Why is Microsoft's agent called Scout instead of Copilot?

Because the architecture changed, not for marketing. Copilot is a promise about assistance: it waits for you to ask. Scout runs on its own schedule, and critically it gets its own Entra identity, acting as a distinct principal with its own permissions inside a zero-trust runtime. An agent that is a separate principal in your directory cannot be a feature of your Copilot session; it needs a separate noun, a separate audit row, and a separate answer for your CISO. Microsoft named that new category Autopilots.

https://falkster.com/answers/why-microsoft-scout-not-copilot

### How do you level and promote PMs when the deliverables are gone?

You level on judgment under uncertainty instead of volume of output. At every level the evidence is a shipped surface with a metric and your fingerprints on why it moved, an eval you own that caught a real failure, and a bet made with the reasoning recorded before the outcome was known. What scales is the size of the bet you are trusted to make without a net: L4 owns one surface and proves the loop works, L5 sets the eval bar others inherit, L6 is trusted with bets that are expensive to reverse, Principal sets the standard the pillar is judged against. The hard part is not the rubric, it is building a calibration room that can tell a good bet from a lucky one.

https://falkster.com/answers/how-do-you-level-pms-without-deliverables

### How can a prototype beat a spec in a product decision?

A spec is a bet placed with no information. A prototype buys the information. When a feature was stuck in a three-week spec review, I built a clickable version in an afternoon, put it in front of five customers, and four of five reached for the wrong control first. That failure mode was invisible in the doc because everyone reviewing it already knew where the button was. The prototype did not win by being better looking. It won by producing evidence, and when the cost of being wrong drops to half a day, evidence beats argument every time.

https://falkster.com/answers/how-does-a-prototype-beat-a-spec

### What are malleable loops in software?

Malleable loops is Dave Killeen's term for improvement moving through the software itself, with almost no homework for the person who spotted the problem or the person it routes to. In his open-source AI Chief of Staff, Dex, the concrete version is two features: Proactive Health, where the agent audits its own automations and connections at session start and repairs what it can, and Dex-to-Dex reporting, where a user's agent writes a privacy-scrubbed defect report, shows the user exactly what leaves their machine, and sends it to the maintainer's agent, which ships the fix and closes the loop with a thank-you. The loop matters more than the malleability: personal customization produces private forks, while the loop routes what each user learns back into the shared product.

https://falkster.com/answers/what-are-malleable-loops

### What can an eval catch that a demo cannot?

A demo is a performance on inputs someone chose; an eval is a measurement across inputs that represent reality. We demoed an AI feature on three hand-picked cases and it was flawless, then ran the eval against 200 real cases split into named slices. On average it scored fine. On one slice, about a fifth of real traffic, it produced confidently wrong output, the kind of failure worse than an error because it looks right. The demo was blind to it because a demo never samples the case that breaks the thing. The eval caught it because named slices expose what an average hides. A demo can start a shipping conversation; only an eval can end one.

https://falkster.com/answers/what-can-an-eval-catch-that-a-demo-cannot

### What does it cost to land a product versus build it?

Build cost lives inside your org as engineer-weeks and compute, is highly visible, and is collapsing because building is a technical problem AI is good at. Landing cost lives mostly on the customer's side as switching effort, sustained GTM attention, and change management, is nearly invisible until it fails, and is flat because landing is a human problem no model has made faster. The ratio between them flipped: building is now the rounding error and landing is the mountain. Most orgs still budget as if building were expensive, which is why they fund the build and assume the landing, and get month-two silence.

https://falkster.com/answers/what-does-it-cost-to-land-a-product-vs-build-it

### What is outcome-based pricing for AI agents and how does it work?

Outcome-based pricing charges for the result the customer wanted, not for access or consumption. Sierra's mechanics: if the AI agent resolves a case with no human intervention, the customer pays a pre-negotiated rate, reportedly around $1.50 per resolution; if it escalates to a human, it is free. It differs from usage pricing because tokens do not correlate with business value: an agent that used a hundredth of the tokens but produced a tenth of the sales would be worse, not better. The design makes the vendor eat the cost of its own failures, which forces accountability for the last mile of implementation. It breaks where outcomes are fuzzy or gameable, in which case honest usage pricing beats a fake outcome metric.

https://falkster.com/answers/what-is-outcome-based-pricing-for-ai-agents

### What does it mean to land a product, not just ship it?

Landing is when a customer's behavior durably changes and the value sticks. It is distinct from shipping (code in production) and launching (the announcement), which you fully control, because landing happens on the customer's side of the glass. The test is strict: point at a behavior a customer used to do a different way, now does with your product, and would be annoyed to lose. Not signups, not week-one activation, not a good demo. It fails four ways: the demo that never becomes a habit, the onboarding that activates but goes silent by month two, the launch with no owner past GA, and the feature that gets used but moves no number.

https://falkster.com/answers/what-is-product-landing

### Which companies treat product landing as a real discipline?

Snowflake made landing its business model by pricing on consumption: customers pay as they run queries and workloads, so revenue only arrives when the product is actually used. That converts landing from a thing nobody owns into the only way the company gets paid, which shows up as net revenue retention around 126 to 127 percent, meaning existing customers expand materially every year. A signup that never becomes usage is worth zero, so no function can book its win at the contract. The transferable lesson is not to copy the meter but to denominate your revenue, or at least your internal scorecard, in the thing that only happens when you land.

https://falkster.com/answers/which-companies-treat-landing-as-a-discipline

### Who owns product landing in an organization?

In most companies, nobody owns landing, and that is the problem. Product owns shipping and hands off at GA, marketing owns the launch and moves on, sales owns the close, and success owns the renewal three quarters later. The ninety days that decide whether a launch mattered fall into the seam between all four. The org chart was drawn when building was the constraint, so every function is optimized around the moment of release. The weak fix is a Head of Landing with no authority. The strong fixes are structural: give landing to whoever owns the outcome and remove the handoff, or price so revenue only arrives on usage, which makes landing everyone's job because it is the only way anyone gets paid.

https://falkster.com/answers/who-owns-product-landing

### Whose job does the Product Builder shift threaten, and how do you lead them through it?

The Product Builder shift threatens the person whose craft is the thing AI made cheap, often your strongest performer on the old scoreboard: the best spec and strategy writer, whose value and identity were built on producing the exact artifact a model now generates for free. The threat is about scarcity, not competence. The tell is that they work harder at the deflating skill, because you cannot out-effort a change in what is scarce. Lead through it by naming the change early, separating the person's worth from the deflating craft, and offering a real path into the scarce skills. The cruelty is never the honest conversation. It is the delayed one, which steals the runway they need to reskill with dignity.

https://falkster.com/answers/whose-job-does-the-product-builder-shift-threaten

### Why are SaaS revenue multiples collapsing in the AI era?

There are two repricings happening at once, and confusing them is dangerous. One is AI destroying the option on moats denominated in accumulated human effort or content: Airtable sold at about 2.7x ARR, Chegg lost roughly 99% of its value as ChatGPT and Google AI Overviews ate its answer library, and Prosus wrote Stack Overflow down from $1.8 billion toward $564 million as question volume fell more than three quarters. The other is zero-interest-rate growth assets repricing to cash-flow multiples through private-equity buyouts: Squarespace at $7.2 billion and Smartsheet at $8.4 billion. You wait out a rate repricing. You cannot wait out an AI repricing, because the value of the moat itself fell. Ask what your moat is denominated in to know which one you face.

https://falkster.com/answers/why-are-saas-multiples-collapsing-in-the-ai-era

### Why do successful launches fail to land?

Launch metrics measure the moment of release, not the ninety days that decide whether release mattered. A launch can go green on press, signups, and week-one activation and still flatline two months later, because signing up is not the same as changing a behavior and keeping it. It falls through in three specific places: no owner past general availability, onboarding that stops at a single activation event, and a product that delivers value once and gives no reason to return. Each gap is invisible on launch day because the launch dashboard is built to light up at release and has no way to show month-two silence. Name a landing owner before launch, instrument the curve past week one, and design a reason to come back.

https://falkster.com/answers/why-do-successful-launches-fail-to-land

### Why does Whatnot run product with so few PMs?

Whatnot's product org was founded on the premise 'we regret that product management exists,' which operates as a forcing function: no PM is ever hired by default, only where a specific problem needs one. Roughly 22 senior PMs cover the $8 billion GMV marketplace, one hire from 31,832 applicants over two years. Every six months leadership stack-ranks the critical problems and assigns each a named DRI, who can be a PM, engineer, or designer. PM managers spend over 90% of their time on IC work. CPO Tom Verrilli's argument: the industry's pod ratio, one PM per six engineers, infantilized engineers and designers and produced PMs whose specialty was politics. AI removed the places that theater could hide.

https://falkster.com/answers/why-does-whatnot-run-product-with-so-few-pms

### Why is Sierra AI growing so fast?

Sierra went from its February 2024 launch to $100M ARR in seven quarters, $150M in eight, and roughly $200M by May 2026, with 40% of the Fortune 50 as customers and a $15.8B valuation. The growth is an operating-model advantage, not a model advantage: outcome-based pricing where a resolved case bills around $1.50 and an escalation to a human is free, forward-deployed agent engineers who own the last mile of implementation, an engineering discipline where annotated real conversations become regression tests, a land-one-channel-then-expand motion aimed at the biggest brands first, and marketing built on benchmarks and engineering essays rather than slogans. Each choice puts Sierra on the hook for whether the product actually lands.

https://falkster.com/answers/why-is-sierra-growing-so-fast

### Why is the PM career ladder breaking in the AI era?

The PM career ladder runs on documents, but not because documents are the work. Documents are the evidence promotions are built from. In a calibration room, someone points at the PRD or strategy you authored and argues it proves the scope of your next level. AI collapses that artifact from three weeks to an afternoon, which removes the object the ladder was calibrated to measure. The evidence gets better and less legible at the same time: a prototype that settled a debate looks like luck, not seniority. The ladder breaks because its currency, legible output, is being devalued, and outcomes and judgment, harder to see and harder to game, have to replace it.

https://falkster.com/answers/why-is-the-pm-career-ladder-breaking

### Why did Airtable sell for only 2.7x ARR?

Because the multiple was never about the revenue, it was about the option. Airtable sold to Bending Spoons at a $1.285B enterprise value on ~$480M ARR, roughly 2.7x, down from a 40x-forward 2021 raise at ~$11.7B. AI did not take the revenue, which is still sticky and growing 20%. It took the option that Airtable becomes the next platform layer, because the moat was effort-based switching cost, and effort-denominated lock-in deflates as agents make effort cheap. Buyers pay 2.7x for sticky cash flow and 40x for a platform option, and the option is gone.

https://falkster.com/answers/why-did-airtable-sell-for-2-7x-arr

### If building is cheap now, what is the real constraint on product builders?

Landing, not building. The 2026 CPO Insights Report shows speed to market became the number one internal challenge CPOs report, up from 14% to 22% in a year, and says the bottleneck moved from building to launching. Twelve builders shipping twelve products is only a win if twelve products can land. The new scarce skills are judgment about what should exist, evals written before code, downside sized by confidence and reversibility, and adoption. None of them come with a title change.

https://falkster.com/answers/if-building-is-cheap-what-is-the-real-constraint

### Should you build your own data governance for AI agents?

No, unless governance is already your product. That moat took Alation fourteen years and five straight years of Gartner Magic Quadrant leadership, and enterprises already own one catalog or three. Plug into the stack they have instead, which is exactly the opening Alation left when their AIOS FAQ told customers to build agents wherever they like. Design three things now because retrofitting them shows: an audit trail a compliance officer would accept, permissions that travel with the data, and a root-cause split between bad data, missing context, and a broken agent.

https://falkster.com/answers/should-you-build-your-own-data-governance-for-ai-agents

### What is the difference between an AI agent, a workflow, and an automation?

Automations follow fixed rules with no judgment. AI workflows add a language model at specific steps inside a deterministic pipeline, making steps smarter without changing the overall flow. Real agents decide what to do next based on context, adapt to unexpected input, and produce variable outputs. Most things marketed as agents today are actually workflows with a chatbot bolted on.

https://falkster.com/answers/agent-vs-workflow-vs-automation

### Are OKRs still worth it for AI products?

Mostly no, not in their standard quarterly form. The 12-week OKR cycle was built when build cycles were weeks and learning cycles were months. For AI products, a Product Builder ships a prototype in four hours and customer signal arrives daily, so a quarterly commit is a lagging response. Replace OKRs with a Rolling Outcome Ledger: 3 to 7 active outcome bets, each with a hypothesis, a test, a 2 to 4 week decision deadline, and one owner. Keep quarterly cadence only for board narrative and long-lead capital allocation.

https://falkster.com/answers/are-okrs-worth-it-for-ai-products

### Are prompt logs the new switch interview?

For half of what switch interviews do, yes. When a customer types a request into an AI feature, they narrate the job in their own words at the exact moment they wanted it, across your entire user base, with no recruiting and no recall bias. That beats booking eight interviews to guess at demand. Logs win on what and where: what the customer is trying to do and where the product fails. Interviews still win on why, on emotion, and on context. Use logs for breadth, interviews for depth, in that order.

https://falkster.com/answers/are-prompt-logs-the-new-switch-interview

### Can you run a staff meeting off a live dashboard?

Yes, and it was the best operating change I made that year. For one quarter I banned status updates from my weekly product staff meeting and replaced them with a live dashboard a small set of agents refreshed before each meeting: what shipped, what moved, eval scores and trends, cost per outcome by workflow, and a flagged list of anything off track. Everyone read it beforehand, so the meeting opened at the decisions. It killed thirty minutes of narration nobody acted on and surfaced numbers that used to hide inside confident verbal summaries. The one trap is density: keep it to the eight numbers that change decisions or people stop reading it.

https://falkster.com/answers/can-you-run-a-staff-meeting-off-a-dashboard

### Can you trust your product intuition?

Less than you think. Product intuition is pattern-matching on past experience, valuable only while the world stays similar to the one that trained it. The AI era rotates the patterns faster than any human gut can recalibrate, so the more experienced you are, the more confidently you apply rules from a world that already disappeared. The fix is not dashboards, it is cheap tests. Demote intuition from a decision to a hypothesis, then test it against a real customer outcome.

https://falkster.com/answers/can-you-trust-product-intuition

### Do different types of PM need different agent stacks?

Yes. Each of the ten common PM roles needs a different AI agent stack, because the same agent watches completely different signals depending on the job. A Growth PM runs Product Health tuned for DAU/MAU and activation. A Platform PM runs it for API uptime and error rates by endpoint. The Product Health agent shows up in almost every stack, but configured differently each time, which is why configuration matters more than capability. Pick the three agents that match your actual role, run them reliably, then add one at a time.

https://falkster.com/answers/do-different-pms-need-different-agents

### Does AI delete the director of product?

AI does not hit the IC PM or the VP first. It hits the layer in between: the Group PM and Director of Product whose main function is moving information up and down, compressing IC status into leadership summaries and relaying priorities back. That routing work is the most automatable knowledge work there is. The Director role is not deleted, it is split. The routing half gets automated. The coaching-and-deciding half becomes the whole job.

https://falkster.com/answers/does-ai-delete-the-director-of-product

### Does continuous discovery scale for AI-native products?

No. Teresa Torres' continuous discovery is the right model for human-centric SaaS and the wrong one for AI-native products, because agent products break the deterministic feedback loop it assumes in three places: the agent generates behaviors users never anticipate, causality is fuzzy in non-deterministic systems, and the weekly cadence is too slow. The alternative is punctuated discovery: three-week intense discovery sprints sandwiched between eight-week build phases, with a 60/40 split of agent testing to customer interviews.

https://falkster.com/answers/does-continuous-discovery-scale-for-ai

### Does jobs-to-be-done still work in the AI era?

Yes, but with a clock attached. Jobs to Be Done rests on one assumption: the underlying job is stable even as products change. AI breaks it by absorbing the job, so the need your customer hired you for can vanish inside a release cycle. The job now has a half-life, and it is shrinking from years to weeks. JTBD still names the job for a moment in time, but you can no longer interview your way to a job map and reuse it for two years. Measure the decay continuously.

https://falkster.com/answers/does-jobs-to-be-done-still-work

### What should you measure weekly versus quarterly for an AI product?

Use two measurement layers on purpose. On the fast cadence (weekly, and per diff) measure direction with seven leading indicators: eval pass rate, agent quality, iteration count, design coherence, escalation rate, dispute rate, and latency. On the slow cadence (monthly, quarterly) measure outcomes. The reason is timing: AI features iterate ten to twenty times a week, but an outcome cycle still runs four to twelve weeks, so by the time an outcome attributes back you have shipped forty to eighty more changes. Outcome accountability alone becomes a lagging signal that cannot drive daily decisions. Most teams run only one layer; the fix is to run both.

https://falkster.com/answers/dual-cadence-measurement

### Can a prototype kill a project before you build it?

Yes. 48 hours before a full-squad kickoff, a rough two-hour Claude Code prototype plus five customer calls revealed our elegant linear workflow was structurally backwards: customers wanted exploration before commitment, collaboration during the process, configuration persistence, and batch application over time. We paused, redesigned in a week, shipped in May instead of April, and hit DAU targets in week one. Four PM hours and 2.5 hours of customer time saved 600 engineering hours building the wrong feature.

https://falkster.com/answers/how-a-2-hour-prototype-killed-a-3-month-project

### How do CPOs win the AI transition without stalling?

Stop treating the AI shift as a soft pivot. Per-seat revenue does not gradually migrate to outcome pricing, it splits into two business systems that fight inside one org chart. The winning move is to set a sunset date on the legacy 18-24 months out, invest all innovation in the agent-native successor, run the legacy for cash, and pre-sell the board on the gross margin trough before it arrives. Run the seven-decision sequence in order. Out of order is the most common reason transitions fail.

https://falkster.com/answers/how-do-cpos-win-the-ai-transition

### How does a PM build financial fluency?

The part you need is small and specific. Know five numbers cold: your area's gross margin, cost per outcome of your main workflow, your two biggest cost drivers and which product decisions move them, the payback or LTV to CAC shape of your segment, and your compute bill trend. Then run one 90-minute working session with finance to rebuild cost per outcome from scratch. Internalize that model choice, caching, and workflow design are margin decisions wearing product clothes. Read Berman and Knight, and skip the certificate mills.

https://falkster.com/answers/how-do-pms-build-financial-fluency

### How does a PM ship their first pull request with AI tools?

Start at the bottom of the four-level ladder: copy and configuration changes, then AI prompts on features you own, then small front-end changes, then telemetry and instrumentation. Do not try to become an engineer. Own the surfaces where you already have the most context. Before your first PR, read the engineering style guide and run the PM PR review skill file, which catches roughly 80 percent of the feedback an engineer would otherwise write.

https://falkster.com/answers/how-do-pms-ship-their-first-pull-request

### How do you build a PM second brain from meeting recordings?

Run a four-step pipeline: record every meeting with a tool like Otter, Fireflies, or Grain, transcribe automatically with Whisper, embed the transcripts into a vector database like ChromaDB, Qdrant, or Weaviate so search works on meaning not keywords, then query in natural language. Five meetings is a curiosity, five hundred is institutional intelligence. The point is the compound effect: the longer you run it, the wider the gap between you and someone relying on memory.

https://falkster.com/answers/how-do-you-build-a-pm-second-brain

### How do you decide whether to cannibalize your own product?

Run four diagnostic questions in order. Can the legacy architecture support the successor's quality bar? Is the legacy customer base the right ICP for the successor? Can the company afford the gross margin trough, which compresses from 78-82% to 58-65% during the ramp? Is the buyer the same person? The answers route the decision to one of three operating modes: sunset, refresh, or split. Most CPOs want the answer to be refresh because it is least disruptive. The honest test is the four questions, not the comfort level.

https://falkster.com/answers/how-do-you-decide-whether-to-cannibalize

### How do you do discovery when your customer is an AI agent?

The canonical discovery playbook assumes a human customer. Agents do not have opinions, get frustrated, or answer interviews, so six methods replace it: agent telemetry as the primary signal, failure-mode interviews with the humans who deployed the agent, capability audits, pattern mining in open source and Discord, an agent benchmark suite, and discovery with end users through their agents. Smallest first step: tag every API call by agent versus human, then run a 30-minute failure-mode interview with the operators behind your three most agent-heavy customers.

https://falkster.com/answers/how-do-you-do-discovery-when-your-customer-is-an-agent

### How do you hire a Builder PM?

Hire a Builder PM by testing the actual skill: can this person ship a working prototype in four hours? The old loop tests case studies, prioritization frameworks, and behavioral polish, all of which coaching and LLMs made easy to fake. The new loop has four rounds: a builder take-home (a real customer transcript turned into a working prototype with a 20-pair eval set, a cost estimate, and a rollback plan), a review session, a live pairing session on a production prompt, and one culture question about a decision you got wrong. Fewer candidates pass. The ones who do are visibly better.

https://falkster.com/answers/how-do-you-hire-a-builder-pm

### How do you lead a product team through layoffs?

The cut is a Tuesday; the job is the next eight weeks. Deciding who goes is brutal but bounded. Keeping the survivors building is the part that actually breaks companies, because a layoff sends a signal to everyone who stayed, and if that signal is fear and chaos they freeze and start looking for the door. Less survives than leaders think: institutional memory leaves, projects lose owners, trust takes a hit, and only what you explicitly protect makes it through. Three moves keep survivors building: name the new strategy within days, cut the projects and not just the people, and be honest about what you do not know.

https://falkster.com/answers/how-do-you-lead-a-product-team-through-layoffs

### How do you move from SaaS to service-as-software as a CPO?

You treat it as a multi-year org rebuild, not a pricing change. Service-as-software sells the work itself, resolved tickets or qualified leads, priced by outcome instead of by seat. Five things change inside the product: telemetry shifts to units of work, quality becomes the product not a differentiator, scope needs contract boundaries, workflows go async, and human escalation becomes a core feature. Pricing flips to per-outcome, per-agent, or consumption. The org chart rebuilds so customer success becomes service delivery and product merges with operations. Budget 18 months. The CPOs who bolt outcome pricing onto a SaaS architecture end up with the worst of both.

https://falkster.com/answers/how-do-you-move-from-saas-to-service-as-software

### How do you run a 15-minute sprint retro that improves things?

Have AI do the 45 minutes of prep the meeting usually wastes. Every Friday at 4 PM, generate a Sprint Health Report from git commits, merged PRs, Jira tickets, Slack messages, and incidents, synthesized by Claude. Monday's retro is 15 minutes: 2 minutes reviewing the report, 5 on the top pattern, 3 on secondary patterns, 5 to celebrate and commit to one action item. The pattern that makes actions stick: Owner will specific action by when. Commit to one per sprint.

https://falkster.com/answers/how-do-you-run-a-15-minute-sprint-retro

### How do you run a CPO listening tour that works?

Stop running a tour and start running a ledger. Ask every interviewee the same nine questions so answers become comparable, and log every answer as a falsifiable claim with five fields: the claim, the source, your confidence, what contradicts it, and what evidence would settle it. The contradictions between functions are the real signal, not noise to average away. Run a 90-minute Friday synthesis to update confidence only where new evidence arrived, and hold all conclusions for 30 days so your information supply stays uncurated.

https://falkster.com/answers/how-do-you-run-a-cpo-listening-tour

### How do you run your first product trio?

A product trio is three people, a PM, a designer, and a tech lead, who own a product area and make meaningful decisions as a unit. It is an operating rhythm, not a meeting. Block two hours weekly: 50 minutes of discovery together with raw notes, 10 minutes writing what you learned, then an hour deciding together with committed ownership. Alignment happens synchronously during discovery instead of in Slack threads half the team never reads. Start with one area, three people, one problem.

https://falkster.com/answers/how-do-you-run-a-product-trio

### How do you run a QBR that forces decisions, not applause?

Nine slides, 30 minutes, deck sent 48 hours early as a standalone pre-read. Slides 1 through 6 make the quarter legible: scoreboard versus plan with the same five numbers every quarter, what we said versus what happened with misses unspun, a decision review that scores calls separately from outcomes, what we killed, customer signal themes with verbatim quotes, and the margin trend. Slides 7 through 9 make the next quarter decidable: a bets table with confidence and reversibility, the anti-bets you are explicitly not doing, and the asks. The meeting spends most of its time on 7 through 9, where the decisions live. A victory-lap QBR is built to be admired; a working one is built to be decided.

https://falkster.com/answers/how-do-you-run-a-qbr-that-forces-decisions

### How do you run a weekly review that keeps you shipping?

Run three ten-minute phases every Monday. Gather: paste last week's data from engineering, metrics, support, and sales into an AI brief prompt. Triage: answer five decision questions about what changed, what is at risk, what blocks you, and what you missed. Communicate: generate a standup opener for your team. The AI does synthesis, you do judgment. Run a four-week trend prompt at month end to catch slow-moving problems that look like noise week to week.

https://falkster.com/answers/how-do-you-run-a-weekly-review-as-a-pm

### How do you train product judgment deliberately?

Judgment is trainable, but only with reps and a scorecard. The scorecard is a decision log with eight fields, two minutes per entry. The reps are written confidence percentages scored quarterly, premortems on big bets, and an explicit one-way versus two-way door call. Most PMs get the reps, hundreds of decisions a year, and skip the scorecard, so experience accumulates without compounding. Score the decision, not the outcome.

https://falkster.com/answers/how-do-you-train-product-judgment

### How do you write an exec update that lands?

Report decisions, not activity, in a fixed format read in 90 seconds: five numbers, three calls, one ask, one kill line. The scoreboard is five numbers with trend arrows, the same five every time: a direction metric, the eval trend, cost per outcome, adoption, and the top customer-signal theme. Three calls you made, each with the evidence and whether it is reversible. One ask framed as a decision with a deadline and a default. And one line on what you killed or declined. Everything else lives below the fold. Bad updates are activity logs; the reader finishes knowing what the team did and not what you decided or need.

https://falkster.com/answers/how-do-you-write-an-exec-update

### How do you write for machines, not just executives?

Write for the three machine audiences that now read PM writing: coding agents that turn specs into implementations, answer engines that decide whether customers find you, and your own agent fleet running off briefs. Swap the prose PRD for an eval set plus a half-page brief, because examples pin down what adjectives never could. Version prompts like code. Then run the cold-read test: hand the artifact to a model with zero context and measure the gap between what comes back and what you meant.

https://falkster.com/answers/how-do-you-write-for-machines

### How do you write OKRs that don't suck?

Most OKRs are disguised task lists. The test for a real one: if your key result were achieved, how would the behavior of your customers, users, or business actually change? If the answer is we will have shipped the thing we planned, it is a task. Real OKRs describe a future state like reduce mobile onboarding from 8 minutes to under 4, not a project plan like redesign mobile navigation. Pick 3 to 5 objectives per function, 3 to 4 key results each, calibrated to a 70 to 80 percent hit rate.

https://falkster.com/answers/how-do-you-write-okrs-that-dont-suck

### How does a decision log train product judgment?

It forces you to record your confidence as a number and what would change your mind before the outcome arrives, then scores the decision quality separately from the outcome at a quarterly calibration review. That separation dismantles resulting, Annie Duke's term for judging a call purely by how it turned out. Over a quarter you learn where your stated confidence and your actual hit rate diverge, which tells you exactly where to trust your own gut.

https://falkster.com/answers/how-does-a-decision-log-train-judgment

### How does AI change the PM-to-engineer ratio?

The PM-to-engineer ratio was never about engineers, it was about how much build output one PM could feed with decisions and direction. The old anchor was about one PM per seven engineers. When AI raises each engineer's output two to three times, the same team generates far more work per PM, so either each PM covers more surface area or the team needs fewer engineers for the same roadmap. Both reduce PMs per unit of output. The 2025 and 2026 PM layoffs are mostly this ratio correction, not a skills purge.

https://falkster.com/answers/how-does-ai-change-the-pm-to-engineer-ratio

### How is the product manager role changing with AI?

The PM job is being reclassified from coordination to building. AI-forward companies hire 'Product Builders' who ship working artifacts instead of documents. Six skills matter now, in priority order: rapid prototyping, customer proximity, AI fluency, outcomes thinking, storytelling and distribution, and end-to-end ownership. Most PMs are still in observation mode. A smaller group builds. A very small group produces output indistinguishable from a small engineering team's, and that gap is growing.

https://falkster.com/answers/how-is-the-pm-role-changing-with-ai

### How many AI agents does a PM actually need?

About seven, one per stage of the PM operating system, not the 30-to-50 agent sprawl most setups drift into. After running a 39-agent fleet across four product orgs, 13 of the 39 were orphaned, never referenced by any other workflow. A seven-agent fleet captures the same compounding effect with a third of the maintenance and zero orphans by design. The leverage of an agent comes from how deeply it integrates into one team ritual, not from how many agents are in the fleet.

https://falkster.com/answers/how-many-ai-agents-does-a-pm-need

### How do you put stakeholder updates on autopilot?

Build an agent that runs Monday morning, pulls data from six systems (GitHub, Jira, analytics, Zendesk, Salesforce, Slack), and drafts three audience-tailored updates: executive summary, team update, and board update. You read the drafts, edit, and send in 10 to 15 minutes instead of 2 to 3 hours. It saves 5 to 6 hours a week across the teams I have set it up for. Start with manual Claude in browser (5-minute setup), move to Zapier plus the API after three weeks. Do not turn on auto-send until you trust the output for 3 to 4 weeks.

https://falkster.com/answers/how-to-automate-stakeholder-updates

### How do you avoid survivorship bias in product management?

Stop drawing conclusions only from the winners. You hear from users who stayed, look at features that survived your current UX, copy competitors who won, and celebrate A/B wins without studying the losers. The countermeasures: talk to churned users through exit surveys and churn interviews, track non-events like drop-off and non-adoption as seriously as events, run feature post-mortems on what flopped, put churn drivers into prioritization alongside power-user requests, and frame the roadmap around outcomes. Abraham Wald's WWII lesson applies: do not armor the planes that came back, armor where the fatal hits were.

https://falkster.com/answers/how-to-avoid-survivorship-bias-in-product

### How do you build a competitive intelligence agent?

Connect an agent to seven sources (competitor changelogs, G2 and Capterra reviews, Salesforce deal notes, Gong call transcripts, support tickets, Slack, industry news) and have it ship one synthesized report every other Monday. The report runs seven sections, from top five insights to positioning recommendations, all readable in ten minutes. The point is not to research better than you can manually. It is to systematize it so nothing slips. I went from 3 hours a week on competitive research to zero. Start with just Gong and Salesforce in week one.

https://falkster.com/answers/how-to-build-a-competitive-intelligence-agent

### How do you build a daily focus agent for PM work?

Build an agent that runs at 7 AM before your first meeting, reads your Google Calendar, Slack, Jira, and email, and posts the three things that actually matter today. Each candidate priority is scored on three axes: urgency (0-40), customer impact (0-35), and executive visibility (0-25). The top three become your day. The report also surfaces invisible blockers, executive commitments coming due, and Slack threads waiting on your decision. Setup is 30 minutes. The point is to stop checking five tools every morning and start the day knowing what to focus on.

https://falkster.com/answers/how-to-build-a-daily-focus-agent

### How do you catch metric drops automatically with an agent?

Build an agent that checks your KPIs hourly, detects movements outside a defined noise band, correlates the drop against recent changes and 48 hours of customer signal, and then ships a clickable prototype of a candidate fix. By the time you read the Slack alert, a working prototype is already live at a preview URL. Three alert tiers (Note, Alert, Incident) match noise to severity. Build a v1 in three weekend evenings starting with one KPI, one collector, one alert.

https://falkster.com/answers/how-to-build-a-kpi-watchdog-agent

### How do you generate launch comms across channels with an agent?

Build an agent that fires when a feature ships, reads its full context (PRD, prototype, Linear ticket, release notes, customer story), and generates every channel of go-to-market copy in one pass: website hero, LinkedIn post, customer email, in-app banner, changelog, and X thread. Six channels in under 4 minutes, all tagged pending review. A voice.yaml file in the repo keeps the tone consistent, and Product Marketing edits, approves, and ships. It fixes the blank page, not the judgment.

https://falkster.com/answers/how-to-build-a-launch-comms-agent

### How do you build a PRD generator agent?

Wire an agent to five inputs: prioritized opportunities, a research repository, your design system, technical documentation, and CRM or analytics data. On a weekly schedule it takes the top 2-3 opportunities and drafts a full PRD for each, grounded in real interview quotes and technical constraints, in eight sections from goals and success metrics through risk and mitigation. The point is not to skip PM thinking. It is to skip the blank page, so a PRD takes 2 hours to refine instead of 10 hours to write.

https://falkster.com/answers/how-to-build-a-prd-generator-agent

### How do you build a product health dashboard agent?

Build a weekly agent that pulls 3 weeks of data from Amplitude or Mixpanel and analyzes four dimensions: feature adoption curves (day 1, 3, 7, 14, 30 vs baseline), cohort retention (this week vs 4 and 8 weeks ago), engagement depth (surface vs regular vs power users), and performance trends from DataDog or Sentry. The point is to distinguish not valuable from not discoverable. A feature with 20 percent adoption but power-user retention 40 percent higher than overall is a discovery problem, not a value problem. This is where quarterly strategy gets shaped.

https://falkster.com/answers/how-to-build-a-product-health-dashboard-agent

### How do you build a working prototype in 60 minutes?

Run five phases against the clock. Minutes 0 to 5, frame the problem: write a two to three sentence problem statement, list the user's steps, name the constraints, and decide what you will not build. Minutes 5 to 15, describe the problem and journey to Claude Code and get a first working version deployed. Minutes 15 to 30, iterate the UX with specific prompts. Minutes 30 to 45, add real data and edge cases. Minutes 45 to 60, deploy, test for bugs, and share the link. You are not building a product in 60 minutes. You are learning what product to build.

https://falkster.com/answers/how-to-build-a-prototype-in-60-minutes

### How do you catch product red flags early with an agent?

Run one agent every morning at 9 AM that scans five sources (Zendesk, PagerDuty, Jira, Slack, Salesforce) and posts a prioritized report using a three-tier model: Critical (must handle today), Warning (monitor, plan response), Info (no action). Each item comes with a recommended next step, an owner, and a timeline, so you go from 'what needs my attention?' to action in 10 minutes instead of 60. Over a quarter it reclaims more than 40 hours of focused time. This is the first agent every PM should deploy.

https://falkster.com/answers/how-to-build-a-red-flag-detection-agent

### How do you build a release readiness agent?

Run an agent every Wednesday that audits five critical inputs: PRD approval status, release notes registration, ETA accuracy, feature flag coverage, and GTM readiness. The output is an 11-section report with a critical-gaps dashboard, a full release inventory, named PM action items, and a go/no-go recommendation. It catches the three things that derail every launch: missing PRDs past deadline, ETAs that silently slip past release date, and features that ship with no release notes entry. By week three of running it, launches are boring.

https://falkster.com/answers/how-to-build-a-release-readiness-agent

### How do you run a retrospective with an agent?

Build an agent that runs every Friday at 5 PM after your retro, reads that day's notes, and compares them against the last 8 retros. It detects recurring themes (the slow-deploys complaint mentioned 5 times that nobody fixed), tracks action-item follow-through, and proposes specific playbook updates. The point is to stop re-discovering the same problems every quarter. Three sprints out, you have a living playbook updated with the patterns your team actually learned.

https://falkster.com/answers/how-to-build-a-retrospective-agent

### How do you build a roadmap progress agent?

Run an agent every weekday at 9 AM that cross-references your roadmap against engineering reality across four systems: Jira or Linear for status, engineering tickets for last activity, GitHub commits for real code activity, and Slack for PM updates. It fires a discrepancy alert when an 'in progress' item has zero commits in 5 days, flags stale tickets and missing engineering tickets, and tracks ETA accuracy. The point is to stop operating on faith and start operating on data. Roadmaps and engineering work live in separate systems and drift weekly.

https://falkster.com/answers/how-to-build-a-roadmap-progress-agent

### How do you build a signal map in your first 30 days as a PM?

A signal map lists the roughly seven doors truth enters a building through, then audits each one on four questions: is it captured, synthesized, routed, and to whom. The seven are the support queue, sales call recordings, churn surveys, product analytics, the customer Slack channel, the engineer who quietly knows everything, and the exec who actually talks to customers. In your first 30 days it beats coffee chats, because it builds a model of the org instead of a contact list.

https://falkster.com/answers/how-to-build-a-signal-map-in-your-first-30-days

### How do you automate sprint planning with an agent?

Build an agent that reads your prioritized opportunities, tech debt backlog, team capacity, and historical velocity from Jira or Linear, then breaks each opportunity into user stories with acceptance criteria, estimates story points from historical patterns, fits stories into the sprint respecting capacity and dependencies, and explains every trade-off. The output is a Jira-ready plan with a confidence score and a list of which stories are most likely to slip. Sprint planning drops from about 4 hours to 30 minutes of review.

https://falkster.com/answers/how-to-build-a-sprint-planning-agent

### How do you turn support tickets into product signal with an agent?

Run an agent every morning at 8 AM that pulls all support tickets from the last 24 hours through three parallel analyses: severity clustering by root cause, segment breakdown (Enterprise, Mid-Market, SMB), and trend detection against 7-day and 30-day baselines. It flags new clusters that did not exist a week ago, segments with 20%+ volume spikes, and any category that jumped over 25% versus baseline. The point is to catch the spike in 'API timeout' tickets on day 2, when it is a 30-minute team fix, not day 6 when a customer escalates.

https://falkster.com/answers/how-to-build-a-support-signal-agent

### How do you prioritize tech debt with an agent?

Build a weekly agent that analyzes the last four sprints of Jira or Linear data, correlates each debt area with its velocity impact, and maps upcoming features to the debt they will touch. It ranks fixes by (impact times urgency) divided by effort and attaches a payback calculation, for example fixing the data layer costs 40 dev-weeks but saves 20 dev-weeks per quarter, so it pays back in two quarters. It tells you the auth refactor affects 5 percent of stories while the legacy data layer affects 40, so you fix the data layer first.

https://falkster.com/answers/how-to-build-a-tech-debt-agent

### How do you build a win/loss analysis agent?

Run an agent every other Tuesday at 10 AM that reads every won and lost deal from the last two weeks. It pulls from Salesforce (deal size, close reason, sales notes), call transcripts, and win/loss interviews, then extracts seven categories: win patterns, loss patterns, feature insights, segment insights, competitive positioning, pricing insights, and three product recommendations. The point is to stop building features nobody asked for and start building the ones that move deals. One run might show: we win enterprise on API reliability and lose mid-market on price, and two fixes would convert 30% of mid-market losses.

https://falkster.com/answers/how-to-build-a-win-loss-analysis-agent

### How do you turn a support ticket into a reviewable PR automatically?

Wire an agent to Zendesk and Sentry so it fires on bugs tagged S2 or higher, then let it reproduce the issue, walk the call graph to the broken code, write the smallest fix, add a regression test, and open a pre-reviewed GitHub PR pointed at the engineer on-call. It runs in about 11 minutes across roughly 100 real runs. A human always reviews and merges. About 75 percent of the PRs merge within 48 hours with only minor edits.

https://falkster.com/answers/how-to-build-an-auto-bugfix-agent

### How do you turn a customer request into a prototype in minutes?

Wire an agent to your signal channels, design system, and source code, then let it draft the feature with Claude Code and deploy a preview URL. When a request lands in Slack, Zendesk, Gong, or Salesforce, the agent ingests the context, cross-references the last 90 days, generates the prototype, opens a GitHub branch, files a Linear ticket, and drafts a Notion doc. End to end it takes about 7 to 10 minutes, so the customer reacts to a clickable thing the same day instead of an eight-week PRD.

https://falkster.com/answers/how-to-build-an-instant-prototype-agent

### How do you build a customer interview synthesis agent?

Point an agent at every interview transcript from the week (Otter, Fireflies, or manual notes) plus segment metadata, and have it produce one synthesis report every Wednesday at 10 AM. The report runs six sections: recurring themes across 3+ interviews, top pains vs. gains, feature mentions, testable hypotheses, contradictions, and segment insights. Signal only emerges across 8+ interviews, and manually reading 12 transcripts to build a theme map takes 6 hours. The agent does it in minutes with customer quotes preserved. Feed in your last 10 transcripts and ask for the top three testable hypotheses.

https://falkster.com/answers/how-to-build-an-interview-synthesis-agent

### How do you build an OKR tracking agent?

Connect an agent to your metrics feeds and run two cadences: daily at 4 PM it scores each key result against the trajectory needed and flags anything off-track, weekly on Friday it predicts confidence. At week 5 of a 13-week quarter you should be roughly 38% of the way to goal; if you are at 25%, it flags it now. The weekly model tells you 'you are at 65% with 4 weeks left, current velocity puts you at 82%, here is what would need to change to hit 100%.' The point is to catch slippage when you can still course-correct, not be surprised at quarter end.

https://falkster.com/answers/how-to-build-an-okr-tracking-agent

### How do you prioritize opportunities with an agent?

Build an agent that reads the outputs of all five DISCOVER agents (Support Signal Processing, NPS/CSAT Analysis, Interview Synthesis, Journey Mapping, Customer Segmentation) and produces one prioritized opportunity stack every Friday at 11 AM. It applies three layers of synthesis: cross-dataset validation (opportunities in 2-plus reports get high confidence), impact estimation (segment size times willingness to pay times churn reduction), and dependency mapping (fixing A unlocks B). The output is OST-ready, so roadmap planning becomes a decision meeting instead of a data-gathering session.

https://falkster.com/answers/how-to-build-an-opportunity-prioritization-agent

### How do you connect Claude to your data sources for PM work?

You connect Claude to your data through Model Context Protocol (MCP), Anthropic's standard for wiring Claude to external tools. Each data source gets an MCP server configured in one JSON file, Claude reads live data from all of them in a single prompt, agents run on a cron schedule, and results post to Slack via incoming webhooks. My production setup at Smartcat connects eight-plus sources (Jira, Slack, Zendesk, Salesforce via Databricks, Gong via Weaviate, Google Calendar, Notion, GitHub, Google Drive). The whole thing takes about two hours, and your first agent runs the next morning.

https://falkster.com/answers/how-to-connect-claude-to-your-data

### How do you defend a margin drop to your CFO?

Run it as three sessions over six weeks, not one pitch. Session 1 walks the trough math and gets agreement on the curve. Session 2 negotiates the comp set for board reporting: benchmark against transition peers, not pure-SaaS, because the same 60 percent gross margin reads as catastrophic against one and on-plan against the other. Session 3 agrees the seven leading indicators that go in every board deck for 24 months. The point is to make the CFO co-own the narrative so the trough reads as expected, not as failure.

https://falkster.com/answers/how-to-defend-a-margin-drop-to-your-cfo

### How do you deprecate a feature?

Deprecate on signal, not politics. Every feature should ship with an explicit kill condition written at launch, like: if weekly active usage stays below 2% of MAU for 8 weeks, the feature is reviewed for deprecation. That turns the death of a feature from a political debate into the execution of a decision already made. Pick one of three tiers, soft sunset, hard sunset with migration, or immediate kill, and send individual notices to the specific users who touched the feature in the last 90 days. Then publicly celebrate what you kill so the culture rewards discipline.

https://falkster.com/answers/how-to-deprecate-a-feature

### How do you kill a feature without breaking trust?

Kill a feature by understanding why it failed before you rip it out, so you learn instead of just cutting. When we killed our most-requested feature, Advanced Scheduling, 247 customer requests had turned into fewer than 50 users out of 4,000. Twenty customer calls revealed they could not articulate a use case; they needed an existing feature they could not find. We renamed it, moved it, and added prompts. Adoption jumped to 600 in six weeks and platform activation lifted 15%. Requests are not requirements. Behavior is.

https://falkster.com/answers/how-to-kill-a-feature

### How do you measure agent reliability without gaming it?

Grade every agent on four metrics, not one. Precision (usable output rate), escalation rate (did it flag uncertainty when it should have), time-to-output (wall-clock from trigger to usable result), and trust trend (are reviewers approving more of its work untouched over time). Precision alone lies, which is why it is one of four. The escalation rate is the metric nobody tracks and the one that catches the dangerous silent failures. Trend matters more than any single score.

https://falkster.com/answers/how-to-measure-agent-reliability

### How do you measure PM velocity from signal to ship?

Measure PM velocity as the time a piece of work takes to travel through the seven stages of the PM Operating System, from Sense to Amplify, not just engineering throughput inside Build. A cross-cutting agent stamps every active item with its current stage, computes median and P90 time-in-stage across the portfolio, names the weekly bottleneck, and flags stuck items. It is a meta-agent wired into the outputs of your other agents. Teams that run it typically see median cycle time drop 50 to 70 percent in a quarter.

https://falkster.com/answers/how-to-measure-signal-to-ship-cycle-time

### How do you measure the cost of being wrong for an AI feature?

Stop measuring decisions by size and measure them by reversibility. Ask one question: if we are wrong, can we reverse this on Tuesday? Cheap-to-reverse calls (most UI, copy, experiments behind a flag) get shipped now and learned from, with no meeting wasted. Expensive-to-reverse calls (pricing, data models, anything customers build on, anything an agent does autonomously at scale) get real scrutiny and guardrails sized to the blast radius. The cost of being wrong is no longer measured in calendar time. It is measured in how hard the decision is to undo and how many times an agent could execute it before a human notices.

https://falkster.com/answers/how-to-measure-the-cost-of-being-wrong

### How should you price an AI product?

Stop pricing the seat and price the work the seat is no longer doing. Per-seat breaks for AI because cost-to-serve scales with usage, not license count, and companies still on per-seat in 2026 run gross margins 40 points below those on hybrid or outcome-based models. Pick one of four models (hybrid, outcome-based, tiered consumption, pure usage) and choose a value unit the customer can understand before they sign. The litmus test: can they tell you what 100 of your units would do for them?

https://falkster.com/answers/how-to-price-an-ai-product

### How do you run a customer interview that actually works?

Interview technique has not changed, but everything around it has. Before the call, a prep agent pulls the customer's support history, usage patterns, and NPS score so you walk in already past the surface. During the call, real-time transcription frees you to listen fully instead of taking notes. After, a synthesis agent compares the transcript against hundreds of other data points in minutes. Use the prototype interview format: 30 minutes to confirm the signal, go deep on the customer's workaround, and show a working prototype to get a concrete reaction. Three to five interviews done this way outproduce ten done the old way.

https://falkster.com/answers/how-to-run-a-customer-interview

### How do you run a pricing migration?

Run it as a six-quarter sequence, not a switch you flip. Quarter 1 is internal alignment and four mandatory conversations (CFO, CRO, board, lead customer). Quarter 2 goes live with hybrid pricing for new accounts and a rewritten comp plan. Quarters 3 through 5 migrate customers in three waves, strategic accounts first. Quarter 6 is sunset and the new normal. The gross margin trough bottoms at 58 to 65 percent around months 10 to 12 and recovers to 70 to 75 percent by month 24. Pre-sell that curve to the board before it happens.

https://falkster.com/answers/how-to-run-a-pricing-migration

### How do you run continuous discovery?

Continuous discovery used to mean two customer interviews a week and a monthly synthesis session. Now agents ingest every sales call, support ticket, NPS response, and app review your company generates, extract the signal automatically, and hand you a ranked opportunity brief every Monday morning. When a signal is clear, you prototype in hours and put a working thing in front of the customer the same day. The cycle that took four to eight weeks runs same-day when the signal is obvious and inside a week when the problem needs a human in the room.

https://falkster.com/answers/how-to-run-continuous-discovery

### How do you run empowered teams when the CEO doesn't buy in?

You don't ask permission, you build evidence. Act like you already have autonomy: reframe feature requests as outcomes, run discovery silently and bring findings, prototype before asking for roadmap slots, and show before-and-after metrics on every ship. Then repeat for six months. Leadership says yes not because they became enlightened but because the evidence is overwhelming. I have watched this play out at seven companies, and the pattern is always the same: leaders waiting for permission lose, leaders shipping better outcomes win.

https://falkster.com/answers/how-to-run-empowered-teams-without-ceo-buy-in

### How do you sunset a pricing tier?

Sunset a pricing tier as a four-stage process over 18 to 24 months, because you are ending a revenue line, not deprecating a feature. Internal preparation runs months 0 to 6, strategic account migration months 6 to 9, mid-market and long-tail migration months 9 to 18, and sunset day plus reorg months 18 to 24. The comp asymmetry is the lever: legacy at 50% of historical comp, successor at 150% of equivalent ACV. Five patterns kill most sunsets, so audit those first.

https://falkster.com/answers/how-to-sunset-a-pricing-tier

### How do you test a product assumption in a week?

Identify the riskiest bets under a feature idea, then use one prototype to test them in days. Map assumptions Monday morning, build the prototype Monday afternoon, test with five customers Tuesday and Wednesday, synthesize Thursday, decide Friday. The prototype is the test: instead of designing separate experiments for desirability, usability, and viability, you build one working thing that tests all three at once. Red assumptions, the ones that kill the project if wrong, go first. At Smartcat one prototype test saved eight weeks of building the wrong recommendation system and led to a shipped feature that moved activation by 28 percent.

https://falkster.com/answers/how-to-test-a-product-assumption

### Is Agile dead, and what replaces it?

Agile was a rational response to two 1990s constraints: humans were bad at comprehensive planning, and building was slow. AI collapsed both, so the ceremony economy on top of Agile became overhead that cannot be justified. But what replaces it is not waterfall. It is a spec-first loop: write the architecture document WITH AI in an afternoon, feed it back into AI as primary context, ship a working prototype the same week, validate against real users, and keep the spec as a living artifact. The thing dying is not iteration, it is the planning aversion that got dressed up in Agile language for thirty years.

https://falkster.com/answers/is-agile-dead

### Is being a Product Builder about coding?

No. Product Builder was never about a PM learning to write React. It is about collapsing the distance between deciding and making. The traditional PM job was coordination, a relay race that existed only because turning an idea into a working artifact was expensive and slow. When one person can go from idea to prototype in an afternoon, the specialization that justified the handoffs stops paying for itself. So it is an org-design shift, not a training-budget one. The two skills that matter are prototyping and evaluation.

https://falkster.com/answers/is-being-a-product-builder-about-coding

### Is being AI-first a product decision or a tooling one?

Being AI-first is a product decision, not a tooling one. Using AI means running the process you already have faster: faster specs, faster tickets, faster teardowns, a conveyor belt at higher RPM. Being AI-first means throwing the inherited process out and rebuilding it assuming AI was in the room on day one. Most product orgs are doing the first and calling it the second. The capacity gap is the real story: AI does not shrink the team, it lets the team finally clear a decade of deferred work.

https://falkster.com/answers/is-being-ai-first-a-product-decision

### Is gross margin a product manager's job now?

Yes. In an AI product every feature has a cost line that moves with usage, and the PM is the only person who sees the full customer job end-to-end. The CFO sees a number at month end and the engineer sees a prompt; only you can say a whole flow costs too much and needs rebuilding. AI-first SaaS runs 55 to 70 percent gross margin against traditional SaaS at 78 to 85, and that gap is closed by product choices: model routing, prompt hygiene, caching, and early-exit logic. Watch one metric: cost per successful action by surface.

https://falkster.com/answers/is-gross-margin-a-pm-job

### Is product management dead?

Product management is not dead. It escaped the product team. For twenty years the discipline was confined to R&D because that was the only place building happened. Agents removed the constraint, so finance, legal, HR, sales ops, and RevOps are now structurally product teams designing systems where humans and agents share the work. Aaron Levie called the person doing this an agent operator. The cleaner name is a Product Builder who is also an Agent Operator. The role is multiplying, not shrinking.

https://falkster.com/answers/is-product-management-dead

### Is Scrum obsolete for AI-native teams?

Yes, and the hiring market is voting: Scrum Master roles are being merged into Engineering Manager roles or removed. Scrum borrowed Extreme Programming's cadence and dropped its engineering rigor, and it normalizes failure by reframing missed timelines as learning. The replacement has a name: a four-to-six person builder pod that runs on three artifacts (a prototype, a five-row eval, an outcome ledger) with a 30-minute weekly review and a five-minute Friday update. No sprints, no story points, no Scrum Master.

https://falkster.com/answers/is-scrum-obsolete

### Is spec-driven development (IDSD or SDD) worth it?

No. Spec-Driven Development (SDD) and its rebrand Intent-Driven Software Development (IDSD) are the same methodology layer with a different sticker. Both insert a document between intent and shipped code, both need a vocabulary and templates, and both exist to be sold as consulting and tooling. The intent document is just a PRD that does not admit it is a PRD. The actual fix is to remove the methodology layer: the prototype is the spec, the eval is the acceptance test, and there is nothing in between.

https://falkster.com/answers/is-spec-driven-development-worth-it

### Is taste the last moat in product?

Yes. When execution costs almost nothing, the scarce input becomes judgment: which problem is worth solving, which version is good enough to ship, which feature dazzles in a demo and dies in production. That judgment is taste, and unlike execution it cannot be cleanly delegated, because it is the accumulated pattern recognition of someone who has shipped, been wrong, and felt the cost. Generation got cheap; selection did not. AI did not replace product judgment, it raised the price of not having it. The move is to protect and encode taste, not to hire around it.

https://falkster.com/answers/is-taste-the-last-moat

### Is the PM-as-translator role dead?

Yes, the PM-as-translator role is dead. For fifteen years PMs were professional middleware, reformatting customer input into decks for engineers, waiting on analysts, filing tickets to change a button label. AI and MCP stripped away three layers of translation at once: the planning artifact, the information bottleneck, and the execution handoff. What survives is judgment, customer intuition, cross-functional alignment, and systems thinking. What dies is the planning deck, the weekly metrics meeting, the voice-of-customer monopoly, and the feature-factory PM.

https://falkster.com/answers/is-the-pm-as-translator-role-dead

### Is the PRD dead?

For most feature work, yes. The PRD made sense when building took three months and being wrong was expensive. A PM with AI can now prototype in an hour, so the cost of being wrong dropped from three months of engineering to one afternoon. Four things replace it: a working prototype, a one-page strategic context doc, a 5-minute Loom, and a 30-minute conversation with engineering. Complex features need more prototyping, not more documentation.

https://falkster.com/answers/is-the-prd-dead

### Is the product budget now a compute budget?

Increasingly, yes. In an agent-native company a growing share of product output comes from compute (tokens, tool calls, GPU time) rather than added headcount, so the cost shifts from fixed salaries to variable, usage-driven spend. A team of six running a fleet of agents produces what twenty used to. Three things change for the CPO and CFO: you budget output not heads, marginal cost stops being zero so waste becomes financial, and capacity becomes elastic. The org that plans in this blended unit wins on capital efficiency.

https://falkster.com/answers/is-the-product-budget-a-compute-budget

### What is the difference between a PM and a Product Builder?

A Product Manager wrote specs, ran rituals, and handed work to engineering on a weeks-long cycle. A Product Builder ships working prototypes, writes evals as the contract, owns a surface end-to-end, and runs on a daily eval-driven cadence. The customer-value job is the same. What changed is the medium (document to working artifact) and the cycle time (weeks to days). Customer judgment, prioritization, and taste stay human and matter more.

https://falkster.com/answers/pm-vs-product-builder

### What belongs in a PRD versus a spec versus a brief?

A PRD tried to do three jobs at once: alignment, commitment, and memory. Split them by job. The brief carries the why, the user, the bet, and the one success metric, on a single page you reread. The spec, in the AI era, is not prose at all; the commitments (SLAs, business rules, acceptance criteria) become evals that fail the build. The prototype carries the solution and does the alignment the prose used to attempt. Keep the thinking, drop the twelve-page container.

https://falkster.com/answers/prd-vs-spec-vs-brief

### How do you tell a reversible product bet from an irreversible one?

Ask one question before any bet: is this group stage or knockout? A reversible bet is group stage, a forgiving experiment, pricing test, or engagement feature that can have a bad week and cost almost nothing because there is a next match. An irreversible bet is knockout, the feature where a single silent quality drop churns the account with no second leg to win it back. The common mistake is treating a knockout bet like a group-stage one you can afford to lose. Name the match first, then pick your risk.

https://falkster.com/answers/reversible-vs-irreversible-product-bets

### What is downside exposure and how do you score a feature on it?

Downside exposure answers one question: if this feature's quality quietly dropped 20% tomorrow, how fast would it cost us customers? Score every AI feature on three axes: revenue flowing through the workflow the feature sits in (not revenue attributed to the feature), the cost of a wrong answer (a bad movie pick costs a shrug, a bad dosage extraction costs trust), and reversibility (how hard the trust is to win back, not how hard the bug is to fix). The ranking rarely matches your engagement dashboard, and the top of the list is usually the boring feature buried in the core workflow with the stalest evals. Point your eval budget there.

https://falkster.com/answers/score-downside-exposure

### Should PMs vibe code, and what should they build?

Yes, PMs should learn to vibe code, and they should point it at internal tools and prototypes, never at customer-facing features. Vibe coding is the highest-leverage new skill a PM can add, turning someone who writes specs into someone who builds their own tools in an afternoon. But vibe-coded code optimizes for looking right, and a PM usually cannot tell at the code level whether it is actually right. Build dashboards, feedback-clustering scripts, and throwaway prototypes where the worst case is losing your own time. Do not ship it to customers, where the worst case is a security hole and an engineering team inheriting a liability they did not write.

https://falkster.com/answers/should-pms-vibe-code-internal-tools

### Should you call yourself an AI PM?

No. AI PM, or AI-native PM, describes tool familiarity, not decision ownership, and by 2026 having used AI tools should be true of every PM. The label lets a company signal forward motion in a job posting without answering the harder question of what the PM actually owns now that AI has eaten research synthesis, first-draft specs, and analysis. Once every candidate has the same line on their profile, it stops differentiating anyone. Name the outcome you own, not the tool you use to get there. I own activation tells a hiring manager something they can act on.

https://falkster.com/answers/should-you-call-yourself-an-ai-pm

### Should you kill AI office hours?

Yes. AI office hours are the 2026 version of the agile transformation: a centralized ceremony that signals motion without producing it. I audited the format in seven product orgs in the last year and found high attendance, high sentiment, and zero measurable change in shipping velocity or eval discipline. Replace them with paired shipping sessions, eval reviews, and kill list reviews, all integrated into real work instead of added as new ceremony.

https://falkster.com/answers/should-you-kill-ai-office-hours

### Should you kill the PRD entirely?

For most product work, yes. The PRD was a workaround for expensive engineering time and expensive misunderstanding, and both premises have weakened. Replace it with three artifacts: a working prototype built in Claude Code in four to six hours, a one-page eval rubric with three to five criteria and at least 20 test cases, and a one-page README. Together they cover all five things the PRD used to contain. Done is when the rubric passes the threshold, not when the doc gets approved.

https://falkster.com/answers/should-you-kill-the-prd

### Should you kill the roadmap?

Yes, kill the roadmap as a fixed quarterly commitment, but replace it with something honest, not nothing. The roadmap freezes a plan built on last month's signals and creates a re-plan tax so high that teams execute against plans they know are wrong. I run a one-page live bet portfolio updated every Monday: 5 to 7 active bets, recently killed, and a standing queue. The hardest part is emotional, trading claimed certainty for the real authority of knowing what is actually happening right now.

https://falkster.com/answers/should-you-kill-the-roadmap

### Should you kill the status meeting?

Yes. The status meeting is not a planning practice, it is a trust deficit papered over with calendar time, and the calendar time costs more than fixing the trust deficit would. Replace it with one live product page per product: six auto-updating strips covering product health, adoption, customer signal, active bets, cost, and incidents. At Smartcat this reclaimed about six hours a week of calendar time per PM. The only meetings that survive are the ones where a human decision actually gets made.

https://falkster.com/answers/should-you-kill-the-status-meeting

### Should you merge the PM and product owner roles?

Merging fixes the symptom, not the cause. The PM/PO split is broken, but the Product Owner role exists because Scrum's sprint cadence needs someone to translate strategy into ceremony-shaped artifacts. Merge the roles inside the same Scrum frame and you just hand one person two broken jobs. The real fix is a four-to-six person builder pod that runs on a prototype, a five-row eval, and an outcome ledger. The PO role has no work in that structure.

https://falkster.com/answers/should-you-merge-pm-and-po

### Should you still run an APM program?

No, not in its classic two-year rotational form. LinkedIn CPO Tomer Cohen killed the APM program because it trains for a role the company no longer hires for: PRD writing, stakeholder management, sprint planning, and presentation, most of which an agent now handles. Replace it with a 12-week Product Builder apprenticeship where the apprentice pairs with a senior Product Builder, ships four prototypes, sends one to production, and learns Claude Code and an eval harness on real work. Twelve weeks of intensive shipping builds more useful judgment than two years of rotational coordination.

https://falkster.com/answers/should-you-run-an-apm-program

### What are the three artifacts that replace a PRD?

The PRD splits into three artifacts: a one-page brief that carries the why and the bet and gets reread, a prototype that aligns the team on the solution and is then thrown away, and an eval that encodes the commitments and fails the build when they break. The workflow is five steps: write the brief first, prototype the solution, extract the commitments into evals, build against the evals, and reread the brief at every major decision.

https://falkster.com/answers/three-artifacts-replace-prd

### Were empowered product teams just a ZIRP feature?

Largely, yes. The empowered product team assumes a company can afford to hand cross-functional teams broad autonomy and patient time to discover the right outcome, and that patience was cheap only when capital was cheap. It got hit from both sides: expensive capital killed the patience it required, and AI collapsed the scarce, expensive build capacity the whole apparatus existed to ration. Marty Cagan was not wrong for his era. He was wrong to teach an era-specific configuration as a permanent law. The autonomy survives. The headcount and the ceremony do not.

https://falkster.com/answers/were-empowered-teams-a-zirp-feature

### What are direction metrics?

Direction metrics are leading indicators measured on the cadence of the work itself. For AI-native and agent products, outcome metrics like NRR or CSAT lag four to twelve weeks behind your changes, and with teams iterating ten to twenty times a week, outcomes cannot drive day-to-day decisions. Direction metrics close that gap. The seven that predict outcomes four to eight weeks ahead: eval pass rate, agent quality score, iteration count, design coherence, customer escalation rate, dispute rate on outcome billing, and latency at p95 and p99. You run them on a two-layer system, direction daily and weekly, outcomes monthly and quarterly.

https://falkster.com/answers/what-are-direction-metrics

### What breaks in a 90-day SaaS-to-agents transition?

Three things break in the first 90 days, and none of them show up in the strategy deck. The CFO conversation is not actually done even when you think it is, because the comp set and board narrative tone were never agreed, only the trough math. The lead customer legal review takes 8 weeks, not 4, because the unit definition triggers legal, security, and procurement questions in sequence. And the maintenance team's morale drops sharper than expected because stability is not a story people want to tell. Each is solvable, but together they cost you a quarter of execution speed if they catch you by surprise.

https://falkster.com/answers/what-breaks-in-a-saas-to-agents-transition

### What breaks when you move off per-seat pricing?

Five things break, and the strategy decks never show them. Migration drift on accounts the dashboards call fine but that barely use the product. A cohort you expected to migrate churns instead. The unit definition you wrote turns out too loose at scale. Dispute volume comes in around 12x your projection. And sales reps quietly negotiate unauthorized extensions. At one $40M ARR company this all hit on sunset day, yet by month 22 outcome revenue was 91 percent and blended gross margin recovered to 71 percent. The work is not linear even when the outcome is good.

https://falkster.com/answers/what-breaks-when-you-kill-per-seat-pricing

### What can no product framework teach you?

No framework can teach judgment under ambiguity: knowing which move to ignore, what the data is not telling you, and when to override the method. Frameworks like RICE, JTBD, and opportunity solution trees are training wheels. They teach the moves and give juniors a safe default and a shared language, but balance is earned, not read. It forms only by making consequential decisions and being wrong enough times that the pattern clicks.

https://falkster.com/answers/what-can-no-framework-teach-a-pm

### What Claude skills should every PM build?

Build five Claude Skills first: Customer Interview Synthesis, PRD/One-Pager, Stakeholder Update, Competitive Intel, and OKR/Outcome Tracking. Skills are folders of instructions Claude references automatically, sitting between Projects (focused work) and custom instructions (global). Each takes about 30 minutes to build and pays back in week one. The one trick that makes them reliable is a triggering line in your custom instructions that tells Claude to check your skills folders before responding. Start with the PRD skill today.

https://falkster.com/answers/what-claude-skills-should-pms-build

### What does a $6.5B startup exit actually teach you?

Less than a failure does. My roughly $6.5B exit was driven mostly by market timing and a category rising regardless of what I shipped, so every decision got retroactively stamped genius with no itemized receipt. The startup that returned almost nothing taught me everything, because there was no enormous outcome to launder my judgment through, so I had to examine each call honestly. Success is a laundering machine: it never tells you which input produced the output. The industry has this backwards, worshipping exits and building product wisdom on the handful of teams the market happened to lift. That is survivorship bias, and cheap AI-era building makes it more dangerous.

https://falkster.com/answers/what-did-a-6-5b-exit-teach-you

### What do acquirers actually buy in a startup acquisition?

Almost never the product itself, and rarely the revenue. A strategic acquirer buys one of four things: a capability it cannot build fast enough, a team it wants intact, a defensive block to keep a rival from getting you, or time in a market where being early is the whole game. When I sold MVC to Microsoft, the deal was about a capability and a team, not the standalone product. Founders optimize the wrong asset constantly, polishing the product and revenue while the real thesis is the team or the capability.

https://falkster.com/answers/what-do-acquirers-actually-buy

### What do boards expect from product in 2026?

In 2026 boards expect a CPO to own the quality, cost, and outcome of an agent-augmented product, not just run the roadmap. The new board deck is seven slides over a quarter: an outcome ledger, per-outcome unit economics, an eval scorecard, an agent inventory, cycle time by stage, a headcount-to-output ratio, and a paragraph on what you are willing to be wrong about. Velocity is no longer defensible; per-outcome cost and margin is. If you cannot answer what the eval score is on the agent driving most of your results, you are presenting to a board that has moved past your deck.

https://falkster.com/answers/what-do-boards-expect-from-product-in-2026

### What do founders actually want from a CPO?

Founders do not actually want a CPO. They want their own product judgment cloned and scaled across more surface area than they can personally cover, without losing control of the decisions they care about most. The relationship fails because nobody names this: the CPO arrives expecting real ownership, the founder keeps overriding the calls that matter, and both feel betrayed. The CPOs who succeed absorb the founder's taste before installing their own, earn the cheap decisions first, then expand ownership only after they have earned predictive trust.

https://falkster.com/answers/what-do-founders-want-from-a-cpo

### What does an AI-native CPO do in the first 90 days?

The classic first-90-days plan optimized execution because execution was the constraint. It is not anymore, so the plan shifts. Days 1 to 30, audit reality not the deck: use the product on real tasks, sit in raw customer calls, map what agents build versus people, and find out if anyone reads evals. Days 30 to 60, instrument what matters: cost per outcome by workflow, eval scores with trend lines, a meeting that forces decisions. Days 60 to 90, kill the ceremonies that survive on inertia and defend the few things that are yours: problem selection, the quality bar, the expensive-to-reverse calls.

https://falkster.com/answers/what-does-a-cpo-do-in-the-first-90-days

### What does it mean to ship with observability?

It means no feature leaves staging without the traces, metrics, and evals that will tell you whether it is working, before your first customer hits it. The instrumentation contract is a one-page document written before any code, covering seven items: success metric, leading indicators, cost meter, eval set, trace points, dashboard URL, and kill condition. If any of the seven is missing, the feature is code complete, not done. A feature without observability is a feature you shipped on faith, and faith is not a strategy.

https://falkster.com/answers/what-does-ship-with-observability-mean

### What does 'the eval is the spec' mean?

It means the eval set, not the PRD, is the primary specification for an AI feature. An eval set is 30 to 200 real input/output pairs that define what good looks like, built from actual production inputs, labeled with the correct output, spiked with 10 to 20 adversarial cases, and scored daily against a rubric. Engineering builds against it, and done means the score went up, not that someone approved a document. The eval is the spec, the changelog, and the definition of done all at once.

https://falkster.com/answers/what-does-the-eval-is-the-spec-mean

### What goes in the product section of a board deck?

Seven slides, ten minutes, no feature parade. Slide 1 is the scoreboard: the same five numbers every quarter, a direction metric, the eval trend, margin or cost per outcome, one adoption number, and one retention-relevant signal. Slide 2 is what changed, including the bad news. Slide 3 is the bets table with status, confidence, and a call-it date. Slide 4 is the kill list with redeployment math. Slide 5 is the one risk that matters with its premortem result. Slide 6 is the asks, framed as decisions with dates. Slide 7 is a footnoted appendix. A board allocates capital, blesses hard calls, and opens doors; it cannot act on a screenshot.

https://falkster.com/answers/what-goes-in-a-board-deck-product-section

### What goes in a CPO 30/60/90 plan?

Build it around four audits and three artifacts. Days 1 to 30, run the reality, money, quality, and decision audits while logging every commitment in a trust ledger. Days 31 to 60, instrument what the audits exposed, build a coalition map, and ship one visible decision that signals the new bar. Days 61 to 90, write the kill list, place two or three evidence-backed bets, name the anti-bets, and deliver a day-90 readout structured as an SCQA memo. Do not reorg, do not rewrite strategy in week two.

https://falkster.com/answers/what-goes-in-a-cpo-30-60-90

### What happens to your product after an acquisition?

Frequently the product you built gets killed as a standalone, and that is often the correct outcome, not a tragedy. When the acquirer's distribution matters more than your roadmap, keeping the product alive separately just splits focus. Microsoft killed the standalone product I built after acquiring MVC and absorbed the capability into a much larger platform. The capability lived and reached vastly more customers. The honest test is simple: after the kill, did the value reach more customers or fewer? If more, the kill was right.

https://falkster.com/answers/what-happens-after-your-product-is-acquired

### What happens when you hold AI agents to a performance review?

When you review agents like team members, you find the ones that are quietly making you slower. I run a quarterly review on every agent using four metrics: precision, escalation rate, time-to-output, and trust trend. Last cycle, three failed and I shut them down, not because they never worked, but because they were confidently wrong in ways I could not catch at review time. The reframe is that an agent doing real work is a worker you hired, with a track record you can coach or fire, not magic that either works or does not.

https://falkster.com/answers/what-happens-when-you-review-ai-agents

### What is a Kill List and how do you run one?

A Kill List is the standing counterpart to a roadmap: a maintained, revisited list of features, projects, and rituals that are candidates to stop. You run it on a recurring cadence with one entry test applied to everything already shipped or in flight: would we start this today, knowing what we now know? If the honest answer is no, it goes on the list, and the default becomes killing it unless someone can defend keeping it. The list exists because every incentive rewards launching and punishes stopping, so a good kill needs a structural home or sunk cost wins by default. When building collapses toward free, deciding what not to ship becomes the scarce skill, which makes the Kill List more valuable, not less.

https://falkster.com/answers/what-is-a-kill-list

### What is a one-page brief and what goes in it?

A one-page brief is a single-page product document that replaces the PRD. It has six sections: problem statement, evidence (quotes plus data), proposed solution with a prototype link, success metrics, risks and mitigations, and the ask. The page constraint forces clarity a 20-to-30 page PRD never does. It takes about five minutes to read instead of an hour, and it stays true mid-project instead of going stale by day five.

https://falkster.com/answers/what-is-a-one-page-brief

### What is agent-to-agent dispatch?

Agent-to-agent dispatch is an architecture where AI agents hand work to other agents without a human relaying context in between. A listening agent detects a customer outcome and passes a structured brief straight to a specialist agent, like a prototyping agent that builds the same day. The PM stops being the relay who carries messages between tools and becomes the dispatcher who sets routing policy and owns the quality gates. The control is the eval bar, not the ticket queue.

https://falkster.com/answers/what-is-agent-to-agent-dispatch

### What is an agent operator team?

An agent operator team is the human-in-the-loop that supervises the AI agents doing the work: watching dashboards, filing alerts, slowing review when output looks shaky, and staying attentive without getting distracted. At falkster.ai the team is six pets (three dogs, two cats, nine chickens), which is half satire and half a real point. An all-agent stack needs supervision, not because agents are unreliable, but because nobody in the building is bored enough to watch them. The role puts attention back in the loop without spending salary or attention budget on it.

https://falkster.com/answers/what-is-an-agent-operator-team

### What is an AI product engineer?

An AI product engineer is one person who combines strategic thinking (the PM skill), experience design (the design skill), and building and shipping (the engineering skill), with AI tools amplifying each. It is not a new job title but a description of how the best product people already work, moving fast across the full lifecycle. Things that used to take three people and four weeks now take one person and four days. PMs are well positioned because they already have the hardest skill: customer empathy and strategic thinking.

https://falkster.com/answers/what-is-an-ai-product-engineer

### What is an opportunity solution tree?

An Opportunity Solution Tree connects a business outcome at the top to the customer problems that drive it, the solutions you could build, and the experiments you run to learn what works. Teresa Torres created the framework. Its power is that it separates three thinking modes most PMs conflate: discovery (what are the problems), solution (what could we build), and validation (which solution moves the metric). The structure has not changed. What changed is speed: AI can populate the opportunity layer from hundreds of support tickets, sales calls, and NPS responses in minutes, where interviews used to take weeks.

https://falkster.com/answers/what-is-an-opportunity-solution-tree

### What is dual transformation for a product org?

Dual transformation is running a legacy SaaS product and its agent-native successor as two products on two clocks inside one org. It is two cadences (a six-week legacy cycle and a one-week successor cycle), three talent categories with hard fences (maintenance 5 to 15 percent, successor 60 to 75 percent, bridge 5 to 10 percent), six CEO scoreboard numbers shown side by side, and three rituals. Trying to run both products on one clock kills the new product first.

https://falkster.com/answers/what-is-dual-transformation

### What is mob prototyping?

Mob prototyping is when the product trio, PM, designer, and engineer, sits together for one day a week and builds a working prototype on one screen with AI coding tools. Mob programming fails for production code because writing code is an individual task. Prototyping is the opposite, because it is a decision-making activity, and decisions are faster when the decision-makers react to the same artifact in real time. By end of day you have something to show a customer the next morning.

https://falkster.com/answers/what-is-mob-prototyping

### What is per-outcome pricing?

Per-outcome pricing ties the price to a unit of work delivered, not access to a tool. Instead of $800 per seat, the customer pays something like $1.75 per AI-resolved support ticket where they did not escalate within 48 hours. Running the math usually reveals the SaaS price was 2x to 3x too low, and the customer is happier paying more because the price is tied to something they can measure. It also makes revenue variable, which is why finance teams get nervous.

https://falkster.com/answers/what-is-per-outcome-pricing

### What does it mean to manage a fleet of agents as PM-as-editor?

PM-as-editor is the skill you need once your agent fleet is running: reading agent output the way a senior PM reads a team member's PRD, cutting what does not serve the purpose, and shipping the 80% version instead of perfecting toward 100%. The framework is a four-tier trust ladder, from ships without review down to agent assists but you author. Every edit is training data for the prompt. Spend 60 seconds after each edit noting what changed and why, then update the prompt weekly.

https://falkster.com/answers/what-is-pm-as-editor

### What is prompt ops?

Prompt ops means giving prompts the same lifecycle as production code: version control, pull request review, automated evals on every change, staged rollout at 1%, 10%, then 100%, monitoring, and one-click rollback. A prompt change without an eval run is a deploy without a test. The five-piece stack takes half a day to set up for one prompt. The PM owns the prompt because it is the spec; engineering owns the wiring.

https://falkster.com/answers/what-is-prompt-ops

### What is service-as-software?

Service-as-software is the shift from software that helps people do work to software that does the work itself. In traditional SaaS you give people a tool and they do the work. In service-as-software, AI agents do the work autonomously and humans review, approve, and handle edge cases. The software is not a tool anymore, it is a worker. This changes UX (from input to oversight), reliability, pricing (from per-seat to per-outcome), and the PM role (from tool designer to workforce architect).

https://falkster.com/answers/what-is-service-as-software

### What is substrate-first engineering?

Substrate-first engineering is the idea that in an AI-native org, engineering's highest-leverage work is the substrate that lets PMs and designers ship to production safely, not the features themselves. The substrate is four things: scaffolded environments, guardrails, an eval harness, and isolated deploys. When it is good, a hundred people can build and only the survivors reach production. When it is missing, every prototype is a risk and engineering becomes the queue everything waits behind.

https://falkster.com/answers/what-is-substrate-first-engineering

### What is the AI noise tax?

The AI noise tax is the credibility a product team loses when it performs the AI shift instead of shipping it. Three patterns pay it: LinkedIn superpower posts that wrap two-year-old model behavior in a new badge, SaaS companies publishing agent-strategy essays while shipping a 2019 forms-and-tables UI with a chat sidebar grafted on, and PMs arguing prototyping breaks customer empathy. Each lets a team feel like it is participating in the shift while shipping the same product as last quarter. The fix is the work, not another post.

https://falkster.com/answers/what-is-the-ai-noise-tax

### What is the AI product operating model?

The AI product operating model keeps the pre-AI foundation (empowered teams, continuous discovery, the PM-designer-engineer trio) and rewires the execution around AI agents. Four things change: the PRD dies because a prototype is faster to build than a spec, the trio becomes a quartet or a duo, discovery compresses from four weeks to four days, and execution overhead drops from about 40% of PM time to near zero. The biggest shift: Friday moved from planning the next sprint to shipping the first iteration.

https://falkster.com/answers/what-is-the-ai-product-operating-model

### What is the anti-backlog?

The anti-backlog is what replaces the traditional backlog once you accept that a 400-item Jira backlog is a graveyard pretending to be an inventory. It swaps that graveyard for three things: a live queue capped at two weeks of work and fed only by signal, a hypothesis library for half-formed ideas written as hypotheses, and a kill list of decisions you have already made not to do. None of the three grows unbounded. Migration from a large backlog takes four weeks, and the discomfort peaks in week four when you delete it.

https://falkster.com/answers/what-is-the-anti-backlog

### What is the impact loop?

The Impact Loop is a four-beat operating rhythm that replaces sprints, stand-ups, and roadmap reviews: Sense (know what is happening), Build (make a working response, not a plan for one), Measure (quantify what actually changed), and Amplify (scale what works and kill what does not). It replaces sprints because sprints optimize for predictability and the Impact Loop optimizes for responsiveness. The loop runs continuously, not on a fixed cadence. AI agents handle sensing with a two-minute daily brief, prototyping takes hours, and measurement is automatic and daily. A full loop from signal to validated, profitable change took eight days in the Smartcat example.

https://falkster.com/answers/what-is-the-impact-loop

### What is the Jevons cliff in outcome pricing?

The Jevons cliff is the strategic decision moment when inference costs have dropped enough that your outcome price is no longer at the sweet spot you set it at. Token costs fall roughly 50% per year, so the gap between your price and your cost grows. It arrives as a cliff, not a curve, because a competitor typically undercuts you 12 to 18 months in and forces a reaction. You then have three choices: hold price to capture margin, pass the savings to capture share, or pass 50% and split the difference.

https://falkster.com/answers/what-is-the-jevons-cliff-in-pricing

### What does the org chart for an AI product company look like?

AI coding tools work, but org velocity stays flat because the gains get swallowed by review queues and coordination meetings. The new org chart fixes three layers in order: get every engineer across the AI-adoption line in about 90 days, rebuild code review for AI speed by moving humans to specification and verification, and flatten the coordination layer since the relay work middle managers do is exactly what agents replace. The end state: leadership owns vision and customer signal, small teams of three to five own an outcome, and agents absorb the status-and-scheduling overhead.

https://falkster.com/answers/what-is-the-org-chart-for-an-ai-company

### What is the PM agent stack?

The PM agent stack is a curated set of open-source Claude tools (skills, subagents, MCP servers, hooks, slash commands) layered on Claude Code so the agent can act on your real PM work. It maps 18 tool categories to the 7-stage PM operating system (Sense, Discover, Decide, Build, Ship, Measure, Amplify). It is the bridge you build today while the enterprise-wide AI brain is still 6 to 18 months out in procurement. You can stand up a useful version in a week of evenings, no permission needed.

https://falkster.com/answers/what-is-the-pm-agent-stack

### What does a Product Builder job ladder look like?

A Product Builder ladder has four levels: Product Builder (L4, 4+ years), Senior Product Builder (L5, 6+ years), Staff Product Builder (L6, 9+ years), and Principal Product Builder (L7, 12+ years). Five dimensions scale with level: scope, architecture decisions, mentorship, external voice, and years of experience. The operating model does not scale. At every level the Builder PM ships working prototypes not specs, designs AI systems end-to-end, owns evals, reasons about cost and latency, and validates with users not stakeholders.

https://falkster.com/answers/what-is-the-product-builder-job-ladder

### What is the seven-agent PM fleet?

The seven-agent fleet is the minimum-viable PM agent stack I would build from scratch today, one agent per stage of the PM operating system. The Sentinel (Sense), The Listener (Discover), The Steward (Decide), The Forge (Build), The Ship Brief (Ship), The Compass (Measure), and The Reflector (Amplify). It replaces the 30-to-50-agent sprawl most fleets drift into. The governing rule is that every agent has a named human owner who runs a kill-switch test once a quarter, and any agent nobody misses in a silent week gets retired.

https://falkster.com/answers/what-is-the-seven-agent-reset

### How do three agents collapse an eight-week cycle to ten days?

Three agents (Instant Prototype, Auto Bugfix, and Launch Comms) collapse the signal-to-shipped cycle by killing the handoffs, not by speeding up individual steps. In the composite example, a customer Slack DM at 4:02pm Monday becomes a live, announced feature ten days later, versus eight weeks the old way. Across the portfolio at Smartcat over a quarter, median cycle time dropped about 60%. Humans still make every decision. The agents handle the typing, cross-referencing, and deployment plumbing so people decide faster with better raw material.

https://falkster.com/answers/what-is-the-ten-day-dev-loop

### What five questions make a product review worth the time?

Five questions turn a status update into a decision-forcing review: what did we learn and what did we change because of it, what outcome moved (not what shipped), what is this costing to run, what is the eval score and which way is it trending, and what are we willing to be wrong about. Together they force learning, outcomes, cost, quality, and risk into the open instead of a deck of green progress bars. Most reviews confirm work is happening rather than forcing a decision, which was tolerable when building was slow. Now that building is cheap, are we on track is the least valuable question in the room.

https://falkster.com/answers/what-makes-a-product-review-worth-it

### What replaces the product org by 2028?

By mid-2028 the product org gets smaller, flatter, and eval-centric. Ten checkable predictions: median team size drops from twelve to six, every product hire ships code in their first week, evals become the highest-status and highest-paid skill, per-outcome pricing is mainstream but hybrid, the roadmap becomes a public artifact, PM/engineer/designer titles blur into Product Builder, and APM programs go extinct at top companies. Each prediction ships with a specific test you can run in July 2028.

https://falkster.com/answers/what-replaces-the-product-org-in-2028

### What replaces the product triad and pods in an AI org?

The triad (PM, designer, engineer) and the pod are replaced by two coordinated units: a human department doing judgment work and an agent department doing routine and synthesis work. Pods assumed humans are the unit of execution, and agent fleets break that assumption. A pod of 5 humans plus 8 agents is not a 13-person pod; it is a 5-person team coordinating with an 8-agent department on exception. They connect through a published interface of escalation triggers, dispute mechanisms, and quality SLAs. This changes hiring, career ladders, and operating cadence.

https://falkster.com/answers/what-replaces-the-triad-and-pods

### What should a PM job description look like in 2026?

A 2026 PM job description should lead with what the PM ships in the first 90 days, not what they coordinate. I audited 200 PM job descriptions and 196 could have been written in 2022. A live JD has four tells: a specific prototype expectation in the first 90 days, named tools required (Claude Code, Cursor, an eval harness, a deployment stack), owned metrics rather than influenced ones, and an explicit statement that PRDs are not the output. A dead JD has translation-layer language, PRD-centric duties, stakeholder management as a job function, and no named tools.

https://falkster.com/answers/what-should-a-pm-job-description-say-in-2026

### Which software categories is AI killing?

AI is killing categories that sat on a single thin layer of value: Q&A reference communities, homework help, standalone grammar and spell-check, generic translation, transcription and OCR, resume builders, ATS keyword tools, and generic copywriting. Three forces do the work: search becoming answers, tools becoming agents, and subscriptions becoming outcomes. A second bucket (search, BI, CRM, tier-one support, project management) is reshaping at the core. A short list with real moats (regulated vertical SaaS, identity, hardware-attached software, marketplaces, data plumbing, payment rails) survives.

https://falkster.com/answers/what-software-categories-is-ai-killing

### What should you do instead of a feature request queue?

Replace the feature request queue with three things built around judgment, not build capacity. A signal synthesis layer that clusters all inbound requests into ranked patterns, a prototype-first investigation that builds a one-week prototype for every cluster that hits threshold, and an outcome bet ledger with 3 to 7 active bets and decision deadlines in weeks. The queue rationed scarce engineering time, but with Claude Code a prototype takes four hours, so the bottleneck is now judgment, and you cannot queue judgment. Your 847-item backlog is not an asset, it is debt.

https://falkster.com/answers/what-to-do-instead-of-a-feature-request-queue

### What is wrong with the 'I built it in a weekend' flex?

It celebrates the one thing about building that stopped being hard. Speed to demo is the cheapest and least durable metric there is: with agents, getting to a first working screen is close to free, which is the whole point of vibe coding. A weekend build proves the demo runs, and nothing about whether the thing survives real users, real edge cases, or its own second month. Building fast is a real skill. The flex is the problem, because it sells time to first output as if it were time to durable value. The only question worth asking is whether it was still running the weekend after.

https://falkster.com/answers/whats-wrong-with-the-built-it-in-a-weekend-flex

### When does vibe coding get more expensive than agentic engineering?

At the crossover point, the moment on the cost curve where the rising maintenance cost of vibe coding overtakes the amortized upfront cost of agentic engineering. Google's whitepaper puts vibe coding at 3 to 10x more per feature past that point. Vibe coding is low CapEx and high OpEx, agentic engineering is the reverse. The practical move is to decide whether a thing is a prototype or a product before you build it, so you never let a prototype quietly graduate into production without paying the structure tax.

https://falkster.com/answers/when-does-vibe-coding-get-too-expensive

### When should you not use AI in a product?

Before wrapping any surface in an LLM, walk a five-step decision tree in order and stop at the first yes: can a rule do it, can a query do it, can a form do it, can a heuristic and lookup do it, and only then use a model. The most powerful pattern is the hybrid: use a rule for 80% of cases and escalate to a model for the 20% that need it, which cuts costs in half with no quality drop. The senior move in 2026 is knowing when to take AI out of a flow.

https://falkster.com/answers/when-not-to-use-ai

### Which AI prompts replace the most PM busywork?

The three prompts that save the most hours in my testing are the Jobs-to-Be-Done extractor (pulls structured job statements from interview transcripts), the support-ticket theme clusterer (turns thousands of tickets into a prioritized roadmap input), and the one-pager-from-notes generator (drafts a full spec). Each saves two to three hours on first use. They work because they follow one pattern: give the AI a role plus constraint, paste real artifacts instead of describing them, and demand a specific output structure. Automate the mechanical parts, never the judgment.

https://falkster.com/answers/which-ai-prompts-replace-pm-busywork

### Which mental models break when building gets cheap?

Four of the usual grid invert. Opportunity Cost weakens because the cost of one more prototype approaches zero. Bottleneck misfires because the constraint moved from engineering throughput to taste and problem selection. Sunk Cost Fallacy gets stronger, not weaker, because regenerating is now cheaper than defending. And Local vs Global Optimum survives but loses its excuse, because exploring the global space is finally affordable. The ones never about cost, First Principles, the 5 Whys, and Incentives, get sharper.

https://falkster.com/answers/which-mental-models-break-when-building-gets-cheap

### Why can killing a product beat launching one?

Because a good kill frequently creates more value than a mediocre launch: it frees a team from a losing bet and points them at a winning one. The best decision of my career was a kill, a mostly-built feature I stopped because the honest answer to 'would we start this today' had become no. The industry rewards exactly backwards, because launches are legible and kills are invisible by construction. In an AI-native world where building is nearly free, the scarce skill moves from building things to deciding what not to ship.

https://falkster.com/answers/why-can-a-kill-beat-a-launch

### Why can't users describe what they need?

Users are great at telling you where it hurts but bad at prescribing the medicine. They know their problems deeply because they live with them every day, but the solution they suggest is limited to what they know is possible, which is usually a slightly better version of what already exists. Before the iPhone, people asked for better Blackberry keyboards, not a touchscreen app platform. So treat feature requests as symptoms, not specifications: dig into the why behind every request, prototype the underlying problem's solution, and use feedback as input, not direction.

https://falkster.com/answers/why-cant-users-describe-what-they-need

### Why should you design the tournament instead of picking winners?

Because you cannot prototype all fifty ideas on the backlog, so someone still has to decide, before any evidence exists, which ones earn an hour. When a test cost a quarter of engineering time, the first decision picked the winner. When a test costs an hour, the first decision picks the bracket: which eight of fifty get prototyped at all. Same judgment, radically cheaper to be wrong. Picking winners is a conviction exercise. Designing the tournament is a portfolio exercise. Then let the prototypes settle who wins.

https://falkster.com/answers/why-design-the-tournament-instead-of-picking-winners

### Why do acquired products die inside big companies?

They die in a predictable three-step pattern I watched run three times inside Salesforce. First the team loses its constraint, the scarcity of money, people, and time that forced sharp decisions. Then it inherits the parent's process, the reviews and planning cycles and approval chains built for a giant org. Then it stops shipping, because the velocity and focus that made it worth buying have been engineered out. The cause of death is almost always the same: the parent removed the forcing function and added friction in its place.

https://falkster.com/answers/why-do-acquisitions-die-inside-big-companies

### Why do AI agents fail in production?

Across ten agents I shipped that failed, three patterns dominate. Five failed because of distribution shift, where training data or source documents did not match production. Three failed because the agent's correctness was beside the point, as users, brand, or culture cared about something it was not measuring. Two failed because of cost-compounding autonomy, where an agent got authority to act with no human gate. The common thread: agents got authority before the eval system was load-bearing.

https://falkster.com/answers/why-do-ai-agents-fail

### Why do cheap prototypes sometimes kill good ideas?

Because the room reacts to what is in front of it, and a rough first prototype of a strong idea reads as a weak idea. When building was expensive, the cost forced a hypothesis before anyone built. When building is an afternoon, that friction disappears, so teams substitute the demo for the thinking and bury good concepts they never actually tested. The fix is to write one falsifiable claim before you open the editor.

https://falkster.com/answers/why-do-cheap-prototypes-kill-good-ideas

### Why does distribution beat product in the AI era?

Distribution beats product because it decides how many people ever try the thing and how cheaply you reach the next customer, and AI made that harsher. AI collapsed the cost of building, so product quality gets copied almost immediately and stops being defensible. What does not clone in an afternoon is the channel, the brand, the install base, and the customer relationships. As building got cheap, the moat moved entirely to distribution.

https://falkster.com/answers/why-does-distribution-beat-product

### Why is a reorg a product decision?

Because of Conway's Law: organizations ship their communication structure, so the boundaries between your teams become the seams in your product. A reorg decides which features are cheap and which are structurally impossible, which sets the ceiling on your roadmap before anyone writes a ticket. Most leaders treat reorgs as HR events, headcount and reporting lines, then act surprised when the product mirrors the new boxes. The org chart is the most binding product spec in the building, and almost nobody reviews it as one.

https://falkster.com/answers/why-is-a-reorg-a-product-decision

### Will AI replace product managers?

No, but AI will replace 80% of what most PMs do: writing documents, sitting in meetings, processing information, and coordinating people. That is all information processing, and AI does it faster and cheaper. The PMs who survive are the ones already doing the other 20% well: customer empathy, strategic judgment, cross-functional leadership, and creative vision. The amplified PM uses AI to clear the 80% and gets roughly 4x more time for the work that actually drives products.

https://falkster.com/answers/will-ai-replace-product-managers

### How do you write an eval rubric for an AI feature?

Do not write the rubric first. Grade about thirty real outputs on gut feel with a notes column, then extract the rubric from the disagreements you had with yourself. Keep a dimension only if failing it maps to a real cost (user harm, trust erosion, rework), which is downside exposure applied to the rubric. Score each dimension binary, pass or fail, because a 1-5 scale hides disagreement in the middle. Version the rubric like code, with every model or prompt change as a review trigger, then ship it as a template the team runs on every diff.

https://falkster.com/answers/write-eval-rubric-ai-feature

---

## Going deeper

- Full text of every chapter: https://falkster.com/corpus/falkster-corpus-full.md
- Everything on Enterprise AI Agents: https://falkster.com/corpus/falkster-corpus-enterprise-ai-agents.md
- Everything on AI Product Management: https://falkster.com/corpus/falkster-corpus-ai-product-management.md
- Everything on AI Business Models: https://falkster.com/corpus/falkster-corpus-ai-business-models.md
- Everything on Building a Company: https://falkster.com/corpus/falkster-corpus-building-a-company.md
- Everything on Product Leadership: https://falkster.com/corpus/falkster-corpus-product-leadership.md
- Every URL on the site also serves markdown: append .md to any path, or send Accept: text/markdown
- Index for agents: https://falkster.com/llms.txt and https://falkster.com/agents.md

Built 2026-09-19 from https://falkster.com. Re-download rather than caching: this file is regenerated on publish.
