Topic · 93 pieces

Enterprise AI Agents

Agent count is a vanity metric. What matters is how many deployed agents still complete production work after ninety days, and what each successful outcome costs.

Most writing about enterprise agents is written by people selling them. This is the other kind. Everything here comes out of building and running agents in production, which means most of it is about the parts that break: the ninety-day survival cliff, the intervention rate nobody logs, the difference between an agent that runs and an agent that finishes. If you are being asked how many agents you have deployed, start with the posts on metrics and come back with a better question.

Start here
  1. 01The AI Eval Starter Kit: Six Templates, Error Analysis to GateTeresa Torres's evals method as six copyable templates: error analysis log, type chooser, code assertions, golden dataset, judge prompt, experiment log and gate.4 min read
  2. 02Torres Wrote the Evals Guide PMs Needed. Here Is Step Four.Torres's three-step evals guide is the best hands-on primer for product teams. I'd add a fourth step, a cost column, and one warning about the cheap tier.10 min read
  3. 03The Enforceable HalfSpotify cut agent token costs 90% by routing grunt work to a cheap model. The useful finding is that they could only enforce half their own rule, and why.7 min read

In the handbook

Everything on this topic

93 pieces
  1. 01The AI Eval Starter Kit: Six Templates, Error Analysis to GateTeresa Torres's evals method as six copyable templates: error analysis log, type chooser, code assertions, golden dataset, judge prompt, experiment log and gate.4 min readSep 26
  2. 02Torres Wrote the Evals Guide PMs Needed. Here Is Step Four.Torres's three-step evals guide is the best hands-on primer for product teams. I'd add a fourth step, a cost column, and one warning about the cheap tier.10 min readSep 26
  3. 03The Enforceable HalfSpotify cut agent token costs 90% by routing grunt work to a cheap model. The useful finding is that they could only enforce half their own rule, and why.7 min readSep 26
  4. 04Context Is King. Written Context Is Rented.Everyone is racing for the context layer. The quadrant nobody occupies is empty for a structural reason, and importing your wiki will not fill it.11 min readSep 26
  5. 05The Canvas Was Never the WorkMiro sold at 2.3x ARR while profitable with $435M in net cash. Compression explains the range, not the position. The mechanism is legibility.8 min readSep 26
  6. 06Everyone Wants to Be the Layer UnderneathGlean and Alation arrived at the same posture from opposite ends: be the substrate, not the assistant. You cannot own a daily habit and stay neutral.6 min readAug 26
  7. 07The Wince Is the Spec: Bottom-Up Evals a Model Can't WriteClaude writes half your evals. It's the half that catches nothing. Top-down vs bottom-up evals, and how I turn gut reactions into verifiers on Heidi.7 min readAug 26
  8. 08The IKEA Story Does Not Prove What You Think It ProvesSalesforce fired 4,000, IKEA reskilled 8,500, and LinkedIn decided this settles the AI question. The facts are real. The comparison is junk.9 min readAug 26
  9. 09100 to 3,000 in a Week: Why Microsoft's Best Agent Is Not Called CopilotMicrosoft's Scout agent, formerly ClawPilot, went from 100 to 3,000 users in a week with no mandate. What its naming, identity, and org design teach.10 min readAug 26
  10. 10Malleable Software Was Never the Point. The Loop Is.Dave Killeen's Dex files its own bug reports: your agent tells his agent what broke, the fix ships. Why malleable loops beat malleable software.10 min readAug 26
  11. 11The Eval That Caught What the Demo Did NotThe demo was flawless. The eval I ran against 200 real cases failed on a slice that was a fifth of our traffic. The failure the demo structurally hid.7 min readAug 26
  12. 12The Sierra Playbook: What They Actually Do DifferentlySierra hit $100M ARR in seven quarters and $200M by mid-2026. Not a model story: six operating choices on pricing, accountability, and go-to-market.21 min readAug 26
  13. 13What Anthropic's Dianne Penn Taught Me About the FrontierWho Dianne Penn is, what she leads at Anthropic, and the four ideas she argues for: evals as the new PRD, the jagged edge, product overhang, and judgment.14 min readJul 26
  14. 14The Agent Layer Is Not the Catalog. Enterprises Will Learn This the Hard Way.Alation's AIOS rebrand says what most agent builders miss: it is not trying to be the agent platform, it is trying to be what every agent plugs into.6 min readJul 26
  15. 15Writing for Machines Is the New Writing for ExecutivesHalf your readers are now agents. Writing that executes faithfully is its own craft: specs as eval sets, prompt hygiene, GEO basics, and the cold-read test.9 min readJul 26
  16. 16The Exec Update Template: Five Numbers, Three Calls, One AskAn exec update template read in 90 seconds: five-number scoreboard, three decisions with evidence, one ask. Plus the agent prompt that drafts it.5 min readJul 26
  17. 17The PM First-90 Kit: Contract, Map, and Day-90 Note TemplatesFour templates for a PM's first 90 days: a one-page manager contract, signal map, decision archaeology worksheet, and the day-90 note. Fill in 10 minutes.6 min readJun 26
  18. 18AI Isn't Replacing Developers. It's Eroding Them.Michael Lawrence's piece on AI eroding developer craft lands. The 2020 dev would have caught the null check in 12 seconds. The fix is evals, not nostalgia.8 min readMay 26
  19. 19IDSD Is Spec-Driven Development With a New Acronym. Kill Both.Intent-Driven Software Development (IDSD) is pitched as the fix for SDD. Same methodology layer, different sticker. The prototype is the spec.7 min readMay 26
  20. 20The 5 PM Websites in Your Bookmarks Need to Become Agents.ProductHunt, Growth.Design, Medium, IndieHackers, TechCrunch. Ankush Panday's daily PM list. In 2026 you don't visit them. Five agents watch them for you.7 min readMay 26
  21. 21The Five-Row Eval Template That Replaced My PRDFive-row eval template that replaces the PRD: Behavior, Input, Expected, Scorer, Threshold. Six failure modes that kill bad evals. Worked example included.9 min readMay 26
  22. 22The Builder Trap. Field Notes from Two AI Calls This Week.Two calls this week. A CMO buying more sophisticated silos. An engineer building more sophisticated silos. Same trap. What outcome thinking looks like.13 min readMay 26
  23. 23What a PM Actually Does All Week, Now and in 2028An hour-by-hour breakdown of the modern product manager's week, and a side-by-side comparison of every task before and after the agent fleet rewires the job.14 min readMay 26
  24. 24Direction Dashboard AgentCompiles seven leading indicators every morning, predicts outcomes 4-8 weeks ahead, runs a weekly Goodhart audit. Direction at AI-native velocity.6 min readMay 26
  25. 25Field Report: 90 Days Inside a SaaS-to-Agents TransitionWhat we believed, what we did, what broke, what we changed. An anonymized after-action report from the first 90 days of a $30M ARR cannibalization.9 min readMay 26
  26. 26Renewal Risk Agent for Migration CohortsCatches the accounts whose operational dashboard says fine but whose qualitative signals say otherwise. Three-dimension scoring, 90 days before renewal.6 min readMay 26
  27. 27Board Narrative Drafter AgentThe most expensive document in product management compiled while you sleep. Quarterly board updates from data sources, comp set, and last quarter's narrative.6 min readMay 26
  28. 28The AI Noise Tax. Three Patterns Killing Product Credibility.Three AI patterns I'm done with: LinkedIn 'Claude superpowers' posts, SaaS apps with bolted-on AI sidebars, and PMs claiming prototyping kills empathy.28 min readMay 26
  29. 29Pricing Migration Tracker AgentWatches every migration account daily, classifies drift type, and recommends a specific intervention. Catch quiet drifters in week one, not week six.7 min readMay 26
  30. 30Build a Discovery Agent Stack: Continuous Customer ListeningDiscovery agent stack: seven open-source Claude repos turn customer interviews, support tickets, and public reviews into a continuous loop. Install + recipes.10 min readMay 26
  31. 31Build a Measurement Agent Stack: End the Dashboard Hamster WheelMeasurement agent stack: seven open-source Claude repos turn dashboards into a daily outcome digest, a variance detector, and a stakeholder update writer.10 min readMay 26
  32. 32Build a Prototype Agent Stack: PRD to Working Demo in a DayBuild a prototype agent stack: eight open-source Claude repos take a PM from idea to working prototype in a day, with TDD, design, and security review.10 min readMay 26
  33. 33Meet the Agent Operator Team at Falkster.aiSix new hires at falkster.ai. They manage the AI agents that build the product. They are paid strictly in treats. The team has not missed a single anomaly.8 min readMay 26
  34. 34The PM Agent Stack: A Bridge to the Enterprise AI BrainFive-part series for PMs whose company is 6-18 months from an enterprise-wide AI brain. Build the personal PM agent stack today with open-source Claude repos.13 min readMay 26
  35. 3510 AI Agents I Built That Failed. The Honest Retrospective.Ten AI agents that failed in production: auto-approve expenses, sentiment classifiers, autonomous pricing, stale-docs RAG. The lessons that stuck after each.12 min readMay 26
  36. 36What Replaces the Product Org in 2028: Ten PredictionsTen checkable predictions for product orgs by July 2028: median team size six, evals the highest-status skill, hybrid pricing mainstream, APM programs extinct.12 min readMay 26
  37. 3739 PM AI Agents Deployed: What Stuck, What Died, and WhyAn honest accounting of 39 PM AI agents across 4 product orgs in 80 days. Stage skew, cadence patterns, and the failure mode I kept repeating.20 min readApr 26
  38. 38The Agent Sandbox (PM Version): A Complete User ManualThe agent management system I built to run falkster.ai, now public. Nine views, 39 agents, three workflows, a 4-layer brain. See what an AI-native fleet does.23 min readApr 26
  39. 39The Seven-Agent Reset: A Lean PM AI Fleet, One Agent Per StageStarting over: the seven-agent PM AI fleet I'd build today. One per PM OS stage, each with named signals, a single surface, and a quarterly kill-switch test.21 min readApr 26
  40. 40Auto Bugfix Agent: Zendesk Ticket to Reviewable PR in 11 MinutesAn AI agent that reproduces customer-reported bugs, locates the broken code, writes the fix, adds a regression test, and opens a reviewable PR.5 min readApr 26
  41. 41Instant Prototype Agent: Customer Request to Prototype in MinutesAn AI agent that turns customer feature requests into working prototypes. Deploys a preview URL, opens a PR, files a Linear ticket, drafts a Notion doc.5 min readApr 26
  42. 42Launch Comms Agent: Six Channels Generated on Every ReleaseAI agent that generates website copy, LinkedIn post, customer email, in-app banner, changelog, and X thread when a feature ships. Six channels, four minutes.6 min readApr 26
  43. 43Signal-to-Ship Cycle Time Agent: Measure PM Velocity Across 7 StagesAn AI agent that tracks PM cycle time across the 7 stages of the Product Operating System (Sense to Amplify) and surfaces the weekly bottleneck.14 min readApr 26
  44. 44The Ten-Day Dev Loop: Three AI Agents Collapsed Our 8-Week CycleHow three AI agents (Instant Prototype, Auto Bugfix, Launch Comms) turned a customer Slack DM into a shipped, announced feature in 10 days instead of 8 weeks.11 min readApr 26
  45. 45KPI Watchdog Agent: Catch Metric Drops and Ship a Fix PrototypeA watchdog agent monitors your KPIs hourly, investigates the root cause of any drop, and ships a working prototype fix before you finish your morning coffee.11 min readApr 26
  46. 46I Gave My AI Agents a Performance Review. Three Got Fired.If an agent does the work of a team member, manage it like one. The scorecard, the coaching loop, and the firing criteria for the agents on your product team.7 min readApr 26
  47. 47Stakeholder Communication AgentGenerate tailored updates for different audiences: exec summaries, board updates, investor reports, team briefings. All from the same data.5 min readApr 26
  48. 48Retrospective Synthesis and Learning AgentExtract learnings from sprint retros automatically. Update playbooks, surface patterns, and drive continuous improvement.5 min readApr 26
  49. 49Win/Loss Analysis AgentBi-weekly analysis of won and lost deals. Extract product insights from sales outcomes. What's winning business, and what's holding you back?6 min readApr 26
  50. 50Feature Adoption Tracking AgentDaily adoption curves for new features. Identify stuck cohorts and recommend interventions before adoption stalls.6 min readApr 26
  51. 51Automated PRD GeneratorConvert prioritized opportunities into PRDs automatically. Drafts based on research context, design specs, and technical requirements.5 min readApr 26
  52. 52Automated Release Documentation AgentAuto-generate release notes, customer comms, and doc updates as features ship. Saves hours on documentation and keeps comms consistent.5 min readApr 26
  53. 53Tech Debt Impact and Prioritization AgentWeekly analysis of tech debt impact on velocity. Which debt items actually matter? What's the priority order to reclaim speed?5 min readApr 26
  54. 54Testable Assumptions Tracker AgentConvert opportunities into testable assumptions. Track validation status weekly. Know which assumptions are holding up your roadmap.4 min readApr 26
  55. 55OKR Progress and Prediction AgentDaily OKR tracking with outcome scoring. Weekly confidence predictions for OKR achievement. See what's on track and what needs intervention.5 min readApr 26
  56. 56Opportunity Prioritization and Synthesis AgentWeekly synthesis of all your DISCOVER agent outputs into a prioritized opportunity stack. OST-ready opportunities with impact estimates and dependencies.5 min readApr 26
  57. 57Automated Sprint Planning AgentConvert prioritized opportunities into user stories and sprint plans every Monday. Accounts for team capacity, tech debt, and dependencies.5 min readApr 26
  58. 58Customer Segmentation AgentWeekly updates to your customer segments and personas. Cohort analysis, segment evolution, and behavioral groupings - automatically maintained.5 min readApr 26
  59. 59Customer Interview Synthesis AgentWeekly automated synthesis of your customer interviews. Key themes, hypotheses, and actionable insights - without the tedious manual work.5 min readApr 26
  60. 60Automated Customer Journey MappingBuild rich customer journey maps bi-weekly from session replays, support patterns, and research data. See exactly where users get stuck.5 min readApr 26
  61. 61Every Type of PM Needs a Different Agent StackTechnical PMs, Growth PMs, Platform PMs, and Product Ops have different jobs. Their AI agent stacks should be different too. Here's what works for each role.8 min readApr 26
  62. 62NPS and CSAT Deep Dive AgentDaily and weekly analysis of your NPS and CSAT scores. Segment breakdowns, feature drivers, and what's actually making your customers happy or frustrated.5 min readApr 26
  63. 63Automate Support Pattern DetectionA daily agent that extracts signals from your support queue: which problems are rising, which segments are hurting, and what real trend lies beneath the noise.4 min readApr 26
  64. 64Claude Skills Every PM Should Build TodaySkills are the most underused feature in Claude. Here's how to build five that genuinely improve your daily PM workflow - with exact prompts and templates.9 min readMar 26
  65. 65Build Your Market Intelligence AgentBi-weekly deep dive into your competitive landscape. Win/loss analysis, market shifts, feature parity, and strategic recommendations so you're never surprised.7 min readMar 26
  66. 66Agent-to-Agent Dispatch: The Product Org Chart Nobody Is DesigningEveryone is building single AI agents for PMs. The real shift is agents handing work to other agents, with the PM as dispatcher. Here is the architecture.7 min readMar 26
  67. 67Build Your Release Checker AgentThe final gate before you ship. This Thursday agent verifies QA, docs, GTM materials, and sign-offs, giving a clear go/no-go for each feature.8 min readMar 26
  68. 68Two Weeks of Agent Tuning: What I LearnedMy first agent reports were 60% noise. After two weeks of calibration, they became the first thing I check every morning. Every adjustment I made and why.8 min readMar 26
  69. 69Build Your Daily Red Flag AgentThe first agent every PM should deploy: a morning scan of support, backlog, Slack, and production that tells you exactly what needs your attention today.7 min readMar 26
  70. 70How I Caught a Churn Signal 3 Weeks EarlyA Red Flag agent caught a 400% support spike on a tier-1 account. Three weeks later they renewed at full terms. Here's exactly what happened.6 min readMar 26
  71. 71Build Your Product Health AgentDaily 4 PM pulse check: engagement, activation, performance, and support trends synthesized into one story. Know exactly where your product stands each day.10 min readMar 26
  72. 72Setup Guide: Connecting Claude to Your Data SourcesStep-by-step guide to wiring Claude into your actual work tools via MCP. Based on my production setup at Smartcat.10 min readMar 26
  73. 73Setup Guide: Self-Hosted Agents with OpenClawFor teams that need on-premise data control. Same agents, same data sources, self-hosted infrastructure.5 min readMar 26
  74. 74Build Your Competitive Intelligence AgentStop manual competitive research. This agent monitors competitors, tracks customer mentions, and delivers a weekly intel report so you always know your market.9 min readMar 26
  75. 75Build Your Team Triage AgentYour #team-product channel is a firehose. This agent reads every message, categorizes issues, assigns owners, and surfaces what's unanswered, twice daily.8 min readMar 26
  76. 76Build Your Engineering Capacity AgentEngineering capacity agent: flags overloaded engineers, catches PTO gaps weeks early, and tells you if the roadmap won't fit before the sprint starts.11 min readMar 26
  77. 77Build Your Documentation Gap AgentFeatures ship without docs and customers find the gaps. This agent scans tickets and the release calendar to catch documentation gaps before they become fires.7 min readMar 26
  78. 78Build Your Weekly Ops Digest AgentDaily agents catch fires. This weekly digest spots the trends: recurring issues, degrading resolution times, and systemic problems that daily reports miss.8 min readMar 26
  79. 79Build Your Product Ops AgentDaily product ops report ranked by ARR at risk: top complaints, feature gaps, usage anomalies, and at-risk customers synthesized from five data streams.10 min readMar 26
  80. 80Build Your Daily Focus AgentYour AI chief of staff: reads your calendar, scans Slack, checks the roadmap, and delivers the 3 things that actually matter today before your first meeting.8 min readFeb 26
  81. 81Build Your Customer Commitment AgentSales promised a feature by Q2. CS said 'next month.' This agent tracks every commitment, cross-references the roadmap, and flags what is overdue.10 min readFeb 26
  82. 82Put Stakeholder Updates on Autopilot: The Complete Setup GuideHow to build an AI agent that generates weekly stakeholder updates, project status reports, and exec summaries - so you never write one manually again.18 min readFeb 26
  83. 83Build Your PM Issues AgentCatch slipping features, overdue customer promises, and cross-team dependency failures before they become fires. Scans roadmap and commitments every morning.9 min readFeb 26
  84. 84Build Your Product Health Dashboard AgentGo deeper than daily metrics. Weekly agent analyzing feature adoption curves, cohort retention, engagement depth, and performance trends to shape strategy.8 min readFeb 26
  85. 85Build a Customer Feedback Pipeline in One AfternoonConnect Zendesk, Slack, and app reviews into a single AI-powered insight engine. Step-by-step setup guide with code snippets.9 min readFeb 26
  86. 86Build Your Release Readiness AgentPrevents launch disasters. Checks PRDs, release notes, feature flags, and GTM readiness every Wednesday so nothing ships without proper preparation.10 min readFeb 26
  87. 87Build Your GTM Release Monitoring AgentTrack if features are ready to sell, support, and succeed. Monitors discovery compliance, beta standards, and GTM materials daily across five readiness gates.10 min readFeb 26
  88. 88Set Up a Competitive Intelligence Agent in 30 MinutesBuild an AI agent that monitors your competitors across 8 sources and delivers a weekly brief to Slack. 30-minute setup, no code, runs forever.13 min readFeb 26
  89. 89Build Your Roadmap Progress AgentYour roadmap says in progress but engineering hasn't touched it in weeks. This agent cross-references your backlog against actual dev activity daily.10 min readFeb 26
  90. 90Build Your Executive Report AgentAuto-generated Monday leadership brief: roadmap status, top wins, key risks, and metrics in a 3-minute read. Your CEO walks in already prepared.8 min readFeb 26
  91. 91Your 'AI Agent' Is Probably Just a Cron JobMost things marketed as AI agents are workflows or automations with a chatbot bolted on. Here's how to tell the difference and why it matters for how you build.6 min readDec 24
  92. 92Many AI Agents Are Actually Workflows or Automations in DisguiseAutomations, AI workflows, and real AI agents behave differently. Knowing which you have prevents inflated expectations and missed opportunities.7 min readDec 24
  93. 93AI Agents and the Future of Work: A Pixar-Inspired JourneyA narrative-style exploration of what happens when AI agents become your coworkers - told through a Pixar-inspired story about Casey and the future office.8 min readOct 24
All topicsWorking on this inside a product org? How I help