# The Falkster Corpus: Enterprise AI Agents

How enterprise AI agents actually behave in production: survival rates, cost per outcome, human intervention, and the failure modes nobody puts in the deck.

The claim: Agent count is a vanity metric. What matters is how many deployed agents still complete production work after ninety days, and what each successful outcome costs.

Source: https://falkster.com/corpus · Built 2026-09-19 · Author: Falk Gottlob
Contents: 91 pieces on this topic

## How to use this

Paste this file into your assistant's project knowledge (Claude Projects,
a ChatGPT project, a Cursor rule file, an AGENTS.md), then work normally.
The point is not to ask it about the corpus. The point is that when you ask
it to size a bet, write a brief, or decide what to kill, it answers the way
this practice answers instead of the way the average of the internet answers.

Every entry carries a canonical link. When something here matters to a
decision, follow the link and read the argument. A summary is enough to act
on and not enough to disagree with.

## Attribution

Written by Falk Gottlob. Free to use for your own work and your team's.
When it shows up in something public, cite it as: Falk Gottlob, falkster.com,
with the canonical link. Not licensed for republication, resale, or model
training.

---

## Everything on this topic (91 pieces)

Each entry is the piece's own extractable summary. Follow the link for the argument.

### The Enforceable Half

Published: 2026-09-17
Canonical: https://falkster.com/blog/the-enforceable-half

Spotify's shunt plugin routes an agent's bulk file reads and boilerplate writes to a cheap model and measured a mean 90% token saving on bulk reads, with 10 to 30 seconds of overhead per delegation. The first version put the routing rule in CLAUDE.md and Claude ignored it, so they replaced advice with hooks that block reads over 350 lines. The finding worth keeping is that only the read path could be gated. Writing boilerplate is a judgment call, and a judgment cannot be enforced at a gate without doing the reasoning the gate was built to avoid. The rule under the rule: you can enforce what is observable at a chokepoint and decidable without the agent's reasoning, and everything else is advisory no matter how it is worded. And routing to a cheap tier is a trust decision disguised as a cost decision. Record which tier produced each fact.

### Context Is King. Written Context Is Rented.

Published: 2026-09-14
Canonical: https://falkster.com/blog/written-context-is-rented

Evan Armstrong's [Context is King](https://www.gettheleverage.com/p/context-is-king) is the sharpest piece written on enterprise AI this year, and the chart everyone is sharing comes out of it. His argument holds. Software is splitting into three layers, systems of record and point solutions are both commoditizing, and Christensen's Law says the margin migrates to the layer in between, the one holding organizational meaning and institutional judgment. Where I'd push is on what actually constitutes that layer. Evan writes one sentence about execution traces feeding back to make the next run smarter and files it under "best of all," a nice property of the category. That sentence is the whole category. Everything else being fought over on that chart is a snapshot, and snapshots port. The best objection to the thesis, raised in Evan's own comments, is that companies will just carry their context to whatever vendor is cheapest next quarter. That objection is right about written context and wrong about earned context, and the difference between the two is the entire moat. What doesn't port is the graded record of what agents did inside your tenant, which acts a human reversed, and which rules survived contact with your outcomes. Nobody on that chart has it yet. That's why the box is empty.

### The Canvas Was Never the Work

Published: 2026-09-13
Canonical: https://falkster.com/blog/the-canvas-was-never-the-work

Compression explains the range. It does not explain where inside the range a company lands. Profitable enterprise software with 90% business revenue trades well above 2.3x right now, every week, in deals that get a quarter of this attention. So the interesting question is not why Miro's multiple fell. It is why a company with fundamentals near the top of the cohort cleared at the bottom of it.

### Everyone Wants to Be the Layer Underneath

Published: 2026-08-29
Canonical: https://falkster.com/blog/everyone-wants-to-be-the-layer-underneath

Alation climbed up from the data catalog and told customers to build agents anywhere. Glean stepped back from the assistant and shipped a Model Hub and an MCP Gateway. Opposite ends of the stack, same landing spot: be the substrate, not the assistant. The logic is sound, because the interface layer is where bundling happens and no independent assistant beats "it's already included in your Microsoft license." The problem is that the layer underneath is where Microsoft and Google already live, so you are often retreating into the incumbent's house. And it has a bill: Glean's 45% wDAU/wMAU ratio exists because it owns a surface people open, and plumbing has no such ratio, because nobody opens plumbing. Go fully underneath and you trade a person with a habit for a procurement team with a spreadsheet. The useful question is not which layer to be, it is what has to happen for you to get paid again next year. A human choosing to come back makes you a surface. A system calling you makes you a substrate. What stays scarce is neither the connectors nor the index, it is the permission model and the record of who asked what.

### The Wince Is the Spec: Bottom-Up Evals a Model Can't Write

Published: 2026-08-24
Canonical: https://falkster.com/blog/the-wince-is-the-spec

Top-down evals are the checks you can write from the task description alone: required fields, format, length, no invented facts. A model drafts those well, because they are the brief translated into a new shape, and you should let it. Bottom-up evals come from reading fifty real outputs and noticing what bothers you, and no model can produce them, because they require having read the outputs and caring whether they're good. The conversion rule is one line: every gut reaction becomes a yes/no a grader can answer without judgment, or it gets dropped. On Heidi, that loop caught actions that were permitted, reversible, logged, in budget, and still wrong, which turned into evidence-strength grading and a three-axis autonomy gate. Top-down evals are downstream of the spec. Bottom-up evals are upstream of it, because they're how the spec gets discovered.

### The IKEA Story Does Not Prove What You Think It Proves

Published: 2026-08-16
Canonical: https://falkster.com/blog/ikea-salesforce-ai-story

The viral LinkedIn story, Salesforce cut people while IKEA reskilled them, gets quoted as proof that AI automation fails and human judgment wins. The facts are mostly real, but the comparison is junk and the conclusion is backwards. The two companies used different technology a decade apart: IKEA's Billie was a 2021 intent router, Salesforce's Agentforce a 2025 agent system. The famous 47% is a containment rate misread as a quality score. The load-bearing claims, rehiring at 1.5x and a Benioff regret quote, are unsourced. More important, both halves are stories of AI working: IKEA only freed capacity because Billie succeeded, and both firms ran the same play, moving deflected-work staff toward revenue. The real lesson is not augment instead of replace. It is that freed capacity becomes a new business line only when there is unserved demand next to the work you automated, and almost nobody checks whether theirs does.

### 100 to 3,000 in a Week: Why Microsoft's Best Agent Is Not Called Copilot

Published: 2026-08-15
Canonical: https://falkster.com/blog/microsoft-scout-not-copilot

Microsoft's most-adopted internal agent is not Copilot. It is Scout, formerly ClawPilot, an always-on desktop agent that went from about 100 to more than 3,000 daily users inside Microsoft in a single week, with no mandate and no campaign. The name is not marketing. Scout gets its own Entra identity and runs in a zero-trust runtime, so it is a separate principal in your directory, which forces a separate noun, a separate audit row, and a separate answer for your CISO. Copilot waits for you to ask. Scout acts on a schedule. The lesson for anyone shipping agents: internal pull is the only pre-launch eval that means anything, identity comes before capability, and if you measure an autopilot by daily active users you kill the exact thing that made it spread.

### Malleable Software Was Never the Point. The Loop Is.

Published: 2026-08-12
Canonical: https://falkster.com/blog/malleable-loops

Dave Killeen's Dex shipped two features that he frames as malleable loops: Proactive Health, where the agent audits its own automations and connections at session start and repairs what it can, and Dex-to-Dex reporting, where a user's agent files a privacy-scrubbed defect report with the maintainer's agent and the fix ships with almost zero homework on either side. My opinion: the malleable software framing undersells it. Malleable software has a forty-year graveyard, and even the AI-fixed version, where the agent does the reshaping, ends at a thousand private forks. The loop is the actual invention. Dex-to-Dex is a discovery pipeline that converts user friction into routed, structured, actionable signal, and Proactive Health is a landing instrument that catches month-two silence before the habit dies. Both are stealable today without an open-source chief of staff: give your product a self-check where the user already is, let it file consent-gated failure reports instead of waiting for tickets, and end your docs with prompts an agent can run against the user's own setup. The caveat is that the loop works at n=1 maintainer with total context. At scale it needs triage and trust, or it is a firehose of well-formatted noise.

### The Eval That Caught What the Demo Did Not

Published: 2026-08-12
Canonical: https://falkster.com/blog/the-eval-that-caught-what-the-demo-missed

We built an AI feature, demoed it on three hand-picked cases, and it was flawless. The room wanted to ship. I ran the eval I had written before building: the same feature against 200 real cases, split into named slices. On average it scored well. On one slice, roughly a fifth of our actual traffic, it failed badly, in a way that would have reached customers and embarrassed us. The demo was not dishonest. It was three lucky draws from a distribution with a hole in it, and a demo only ever shows you the draws someone chose. The eval showed the distribution, including the hole. We did not ship. We fixed the slice, re-ran, and shipped a week later on a number instead of a feeling. The demo is a performance. The eval is the only reviewer in the room that is not trying to impress you.

### The Sierra Playbook: What They Actually Do Differently

Published: 2026-08-12
Canonical: https://falkster.com/blog/the-sierra-playbook

Sierra sells AI agents for customer experience, launched February 2024, and hit $100M ARR in seven quarters, faster than almost any software company on record, reaching roughly $200M by May 2026 with 40% of the Fortune 50 as customers. The mechanism is six choices, not one. They price the outcome: a pre-negotiated rate per resolved case, escalations to humans free, which makes Sierra's revenue depend on the product actually working. They own the last mile with forward-deployed agent engineers, deliberately absorbing the implementation risk every SaaS vendor spent two decades pushing onto the customer. They invented an engineering discipline for non-deterministic software, where annotated real conversations become regression tests so an agent never makes the same mistake twice. They build capabilities ahead of the models and throw the code away without sentiment when models catch up. They land one high-volume channel at a giant brand, prove resolution and CSAT, then expand. And they market with credibility artifacts, public benchmarks and engineering essays, instead of category slogans. Each choice is stealable. The compounding of all six is the moat.

### What Anthropic's Dianne Penn Taught Me About the Frontier

Published: 2026-07-24
Canonical: https://falkster.com/blog/anthropic-dianne-penn-takeaways

Dianne Na Penn is Head of Product for AI Research and Labs at Anthropic, the company's first technical product manager, hired in 2023 when there were about five product engineers. Her throughline is that building on the frontier is a different craft than building normal software, and the tools have to change to match. Evals replace PRDs because a probabilistic system needs a runnable definition of success, not a prose description of intent. Capabilities arrive discontinuously, so you need evals as a radar to catch the jumps before they sit unclaimed as overhang. The craft shifts from sweating pixels to reading trajectories and diagnosing why a run failed. Managers who stop building lose their theory of what is possible. And through all of it, human judgment gets more valuable, which is why we will need product people more than ever. My take on all thirteen is below.

### The Agent Layer Is Not the Catalog. Enterprises Will Learn This the Hard Way.

Published: 2026-07-16
Canonical: https://falkster.com/blog/the-agent-layer-is-not-the-catalog

Alation rebranded fourteen years of data catalog work as AIOS, and the tell is in their own FAQ: build your agents wherever you like, Claude Code, n8n, Microsoft 365, and AIOS just makes sure they work from governed data. The company with the deepest enterprise governance moat is not trying to be the agent layer. It is trying to be what every agent layer plugs into. If you are building agentic tools you will not out-govern Alation, Collibra, or Purview, so stop trying and take the opening they are leaving: be the agent platform that plugs into the governance stack a company already has. MCP is the connective tissue that makes that possible. Self-hosted sovereignty is not a privacy nicety, it is the insurance policy compliance officers are shopping for. And the three things worth designing now, because retrofitting them shows, are an audit trail a compliance officer would accept, permissions that travel with the data, and a root-cause split between bad data, missing context, and a broken agent.

### Writing for Machines Is the New Writing for Executives

Published: 2026-07-16
Canonical: https://falkster.com/blog/writing-for-machines

Three machine audiences read PM writing now, and each rewards something different: coding agents reward unambiguous acceptance criteria and labeled examples, answer engines reward extractable, entity-rich blocks, and your own agent fleet rewards stable schemas and brief files. The craft shift is from prose PRD to eval-set-plus-brief, because examples pin down what adjectives never could, a discipline the handbook covers in [The Eval Is the Spec](/handbook/the-eval-is-the-spec). Prompt hygiene means versioning prompts like code, per [Prompt Ops](/handbook/prompt-ops). And the only eval of your writing that matters is the cold-read test: hand the artifact to a model with zero verbal context and measure the gap between what comes back and what you meant.

### The Exec Update Template: Five Numbers, Three Calls, One Ask

Published: 2026-07-07
Canonical: https://falkster.com/blog/exec-update-template

The exec update template has four sections above the fold. A scoreboard of five numbers with trend arrows, the same five every time: direction metric, eval trend, cost per outcome, adoption, and the top customer-signal theme. Three calls you made, each with the evidence and whether it is reversible. One ask, framed as a decision with a deadline and a default. And one line on what you killed or declined. An agent drafts the mechanical 80% from your tools, covered in [the stakeholder update autopilot](/blog/stakeholder-update-autopilot), and a 10-minute human pass adds the judgment. The format is the executive-communication layer of [the skill stack](/handbook/skill-stack) compressed into one page.

### The PM First-90 Kit: Contract, Map, and Day-90 Note Templates

Published: 2026-06-13
Canonical: https://falkster.com/blog/pm-first-90-kit

Four templates, downloadable below. The manager contract is the one-pager you draft in week one: what you own, what good looks like at day 90, how decisions get made, and the silent worries question that surfaces the real job. The signal map scores the seven sources of customer signal on whether they exist, get captured, get synthesized, and get routed. The decision archaeology worksheet reconstructs the last ten decisions to map the real operating model, with an agent prompt that does most of the digging. The bet one-pager and day-90 note close the quarter with a falsifiable bet and a half-page scoreboard. The plan these serve is [the PM 30/60/90](/handbook/pm-30-60-90); the deep dives are [the first-30-days signal map](/blog/pm-first-30-days-signal-map) and [days 61 to 90, own a bet](/blog/pm-days-61-90-own-a-bet).

### AI Isn't Replacing Developers. It's Eroding Them.

Published: 2026-05-28
Canonical: https://falkster.com/blog/ai-isnt-replacing-developers-its-eroding-them

Michael Lawrence's argument that AI is quietly degrading developer craft is correct in mechanism. When the LLM produces code that compiles and passes tests, the human reviewer's pattern recognition has nothing to push against and atrophies in months. The skill rusts fast for the specific patterns the LLM handles reliably. The piece stops short of a workable fix. Slowing down or hand-writing more code is nostalgia and economic pressure will overrule it. The actual fix is to externalize the rigor into the eval suite (a five-row template that catches the null-check class of regression) and to rebuild the pattern-recognition muscle through paired shipping (one driver, one rider, the rider's job is to catch what the model missed). Skill drift is real. The defense is structural.

### IDSD Is Spec-Driven Development With a New Acronym. Kill Both.

Published: 2026-05-28
Canonical: https://falkster.com/blog/idsd-is-sdd-with-a-new-acronym

Spec-Driven Development is a 2025 consulting cycle that turned tools like spec-kit, Kiro, BMAD, Tessl, and Agent OS into a methodology. Kapil Viren Ahuja correctly names SDD as overfit and rigid. His proposed replacement, Intent-Driven Software Development (IDSD), is the same artifact stack with a different sticker. Intent documents are PRDs that do not admit they are PRDs. The actual fix is to remove the methodology layer entirely. The prototype is the spec. The eval is the acceptance test. There is no document between intent and ship.

### The 5 PM Websites in Your Bookmarks Need to Become Agents.

Published: 2026-05-28
Canonical: https://falkster.com/blog/pm-websites-replaced-by-agents

Ankush Panday recommends five daily PM websites: ProductHunt, Growth.Design, Medium, IndieHackers, TechCrunch. The list is reasonable as a 2020 workflow. In 2026 the signal-to-noise ratio at each site is roughly 1:50 and the 30-minute daily ritual converts to almost zero shipped work. The fix is to retire the website-visit habit and run an agent against each source. ProductHunt becomes a launch-triage agent. Growth.Design becomes a case-study summarizer. Medium becomes a synthesizer over a curated writer list. IndieHackers becomes a founder-pattern agent. TechCrunch becomes a funding-and-pivot agent. The stack runs on RSS plus Claude plus a daily digest. Six hours of one-time setup. Saves 100 hours a year and the signal quality is better, not worse.

### The Five-Row Eval Template That Replaced My PRD

Published: 2026-05-27
Canonical: https://falkster.com/blog/five-row-eval-template

The five-row eval template has exactly five columns: Behavior, Input, Expected, Scorer, Threshold. Each row of the spreadsheet is a falsifiable claim about the feature. Twenty to thirty cases is enough for a typical feature scope, not 200. The PM writes the first draft in three to five hours. The engineer adds adversarial inputs. A customer or domain expert pressure-tests the expected column. The threshold is what makes the eval load-bearing, and is the column most teams skip. Six failure modes kill bad evals: vibes scoring, happy-path-only cases, threshold drift, scorer collusion, ceremony writing, and audit-only running. The template is not new. The discipline of writing it well is the work.

### The Builder Trap. Field Notes from Two AI Calls This Week.

Published: 2026-05-18
Canonical: https://falkster.com/blog/the-builder-trap

I had two AI calls this week. One was with a CMO at a financial institution who is being sold a Salesforce Data Cloud transformation. The other was with an engineer in Latin America building a real product on a stack of vector search, fast APIs, and their own agent layer. They sit on opposite ends of the technical spectrum and they were making the same mistake. The CMO was buying more sophisticated silos. The engineer was building more sophisticated silos. The builder trap is the belief that you are supposed to assemble the parts. You are not. You are supposed to own the outcome and let the agent layer decide what parts are needed at all.

### What a PM Actually Does All Week, Now and in 2028

Published: 2026-05-13
Canonical: https://falkster.com/blog/pm-week-before-and-after-agents

A senior PM at a 50 to 500 person company spends about 50 hours a week today. Roughly 43 of those hours are coordination, triage, and artifact production. Only about 7 hours, around 14% of the week, touch a real bet, a real customer, or a real strategic call. The rest is overhead. In 2028 the hours stay the same. The shape inverts. About 30 of those 50 hours move to designing, reviewing, and operating an agent fleet that does the coordination work in the background. The remaining 20 hours go to the things humans are still uniquely good at: strategic bet selection, customer time at depth, prototyping the actual thing, and shipping a working prototype to at least one real customer every single week. The PM job is not getting eliminated. It is getting compressed, then sharpened. The PMs who survive this are the ones who treat agents as a team they manage, and who walk into every Monday with a working artifact in production, not a doc.

### Direction Dashboard Agent

Published: 2026-05-10
Canonical: https://falkster.com/blog/agent-direction-dashboard

The Direction Dashboard agent pulls seven leading indicators every morning at 6 AM, predicts the outcome metrics they imply 4 to 8 weeks ahead, and posts a chart pack to a #direction Slack channel the whole team reads with morning coffee. The seven indicators are eval pass rate, agent quality score, iteration count, design coherence, customer escalation rate, billing dispute rate, and p95/p99 latency. Every Monday the agent runs a Goodhart audit, checking whether each indicator is still correlating with the outcome it's supposed to predict (above 0.7 healthy, 0.5 to 0.7 weakening, below 0.5 dead). The cultural shift is measuring direction at the cadence of the work, not the cadence of outcomes. Start by writing the `direction-indicators.yaml` registry.

### Field Report: 90 Days Inside a SaaS-to-Agents Transition

Published: 2026-05-09
Canonical: https://falkster.com/blog/field-report-saas-to-agents-90-days

Three things broke in the first 90 days of a $30M ARR SaaS-to-agents transition. The CFO conversation had more sub-agreements than I'd inventoried (comp set, leading indicators, board narrative tone all needed explicit agreement, not just the trough math). Lead customer contract legal review took 8 weeks instead of 4 because the unit definition triggered legal, security, and procurement questions in sequence. Maintenance team morale dropped sharper than expected because "stability" isn't a story people want to tell; renaming the team to "Migration Engineering" and tying their bonus to migration tooling adoption fixed it. The strategic frame, operating cadence, board pre-sell, and team belief all held. The first 90 days were harder than modeled and on track for the second 90 to be easier.

### Renewal Risk Agent for Migration Cohorts

Published: 2026-05-08
Canonical: https://falkster.com/blog/agent-renewal-risk-during-migration

The Renewal Risk agent catches the accounts whose operational dashboard says "on plan" but whose qualitative signals say "forming a negative opinion." It scores every migrating account weekly on three dimensions (pricing comfort, product trust, relationship health) using support tickets, Gong sales call transcripts, customer emails, and NPS comments. The high-value signal is divergence: operational fine, qualitative concerning. The agent classifies risk into five types (pricing skepticism, quality concern, competitive evaluation, internal political risk, value misalignment) and recommends a specific intervention. 90 days before renewal, not at the renewal call. By then it's too late.

### Board Narrative Drafter Agent

Published: 2026-05-06
Canonical: https://falkster.com/blog/agent-board-narrative-drafter

The Board Narrative Drafter agent turns the most expensive document in product management, the quarterly board update, into a ninety-minute edit instead of two to four full days. It ingests the quarter's data from billing, CRM, product analytics, the migration tracker, and the eval system, runs it against the CFO-agreed comp set (transition peers, not pure-SaaS), and composes the four-section narrative the board has agreed to. The CPO does the editorial layer: which moments matter, which framing works, what to cut. Start by writing the two YAML files this agent needs, `board-data-sources.yaml` and `comp-set.yaml`. That alone saves a day next quarter.

### The AI Noise Tax. Three Patterns Killing Product Credibility.

Published: 2026-05-06
Canonical: https://falkster.com/blog/ai-noise-tax

Three AI noise patterns are eating product credibility in 2026. First, "Claude superpowers" posts on LinkedIn and Instagram, content that wraps two-year-old model behavior in a new badge every week. Second, SaaS companies publishing thousand-word essays on their AI agent strategy while shipping the same 2019-era forms-and-tables UI with a chat sidebar grafted onto the corner. Third, PMs writing posts arguing that prototyping breaks empathy with customers, defending a workflow that depended on the engineering capacity they no longer have. Each pattern lets a team feel like it is participating in the AI shift while shipping the same product as last quarter. The fix is the work: real changelog evidence, agent-as-surface UIs, prototype-driven discovery.

### Pricing Migration Tracker Agent

Published: 2026-05-04
Canonical: https://falkster.com/blog/agent-pricing-migration-tracker

The Pricing Migration Tracker watches every account on your migration plan daily, scores actual behavior against committed terms, and classifies drift type into one of five buckets: underutilization, quality concern, trust concern, commercial concern, or multi-concern. Each drift type maps to a specific intervention (usage check-in, CS-led dispute review, CPO involvement, CFO escalation) with a 48-hour SLA. The point is the loud-customer model fails because quiet drifters get missed. The agent makes silence visible. The cultural shift is "no news" stops being interpreted as good news. Start by writing the `migrations.yaml` roster by hand. Forty accounts max. That alone surfaces the drift problem.

### Build a Discovery Agent Stack: Continuous Customer Listening

Published: 2026-05-04
Canonical: https://falkster.com/blog/build-discovery-agent-stack

This post is the concrete how-to for the Discover stage of [the PM operating system](/handbook/product-operating-model). It is part 2 of 5 of the PM Agent Stack series. If you have not read [the overview](/blog/pm-agent-stack-overview) yet, start there. It sets up the destination (an enterprise-wide AI brain), the gap (why most companies are 6 to 18 months away), and why the bridge (this stack) is what to build today.

### Build a Measurement Agent Stack: End the Dashboard Hamster Wheel

Published: 2026-05-04
Canonical: https://falkster.com/blog/build-measurement-agent-stack

This post is part 4 of 5 of the PM Agent Stack series. It is the concrete how-to for the Measure stage of [the PM operating system](/handbook/product-operating-model). If you have not read [the overview](/blog/pm-agent-stack-overview) yet, start there. It sets up the destination (an enterprise-wide AI brain), the gap, and why this stack is what to build today.

### Build a Prototype Agent Stack: PRD to Working Demo in a Day

Published: 2026-05-04
Canonical: https://falkster.com/blog/build-prototype-agent-stack

This post is part 3 of 5 of the PM Agent Stack series. It is the concrete how-to for the Build stage of [the PM operating system](/handbook/product-operating-model). If you have not read [the overview](/blog/pm-agent-stack-overview) yet, start there. It sets up the destination (an enterprise-wide AI brain), the gap, and why this stack is what to build today.

### Meet the Agent Operator Team at Falkster.ai

Published: 2026-05-04
Canonical: https://falkster.com/blog/meet-the-agent-operator-team

This week I onboarded the Agent Operator team at falkster.ai. Six new hires. Three species. Median experience level: very alert. Compensation: treats.

### The PM Agent Stack: A Bridge to the Enterprise AI Brain

Published: 2026-05-04
Canonical: https://falkster.com/blog/pm-agent-stack-overview

The destination for every product organization is one AI brain that has read access to every system the company runs. Slack, email, calendar, meetings, documents, source code, dashboards, CRM, design files. All of it. Not parts. All. Most companies are 6 to 18 months from that destination because procurement, security review, and data governance move slowly and PMs are not the buyers for the platform layer.

### 10 AI Agents I Built That Failed. The Honest Retrospective.

Published: 2026-05-04
Canonical: https://falkster.com/blog/ten-failed-ai-agents

Ten AI agents that failed in production, with the cost of each. The auto-approve expense agent that approved $2,400 in dinners-with-customers before finance turned it off. The customer sentiment classifier that scored 89% on held-out data and 61% in production within six weeks because the training distribution missed non-English customers. The autonomous pricing tester that locked in a 14% under-anchored price for two quarters. The internal docs RAG that confidently surfaced 2021 wiki pages alongside 2024 ones. Six more, each with the same structure: what I tried, what broke, what it cost, what I would do differently. The common thread across all ten: agents got authority before the eval system was load-bearing.

### What Replaces the Product Org in 2028: Ten Predictions

Published: 2026-05-01
Canonical: https://falkster.com/blog/product-org-2028

Ten predictions for what product organizations look like in the first half of 2028. Median team size drops from 12 to 6. Every product hire ships code in the first week. Evals become the highest-status, highest-paid skill in product. Per-outcome pricing is mainstream but hybrid. The roadmap becomes a public artifact, not a planning artifact. The PM/engineer/designer titles blur into Product Builder. Customer Success becomes outcome-attainment, not adoption. Quarterly OKRs survive only for board narrative. The Eval Engineer becomes the new Growth PM. The product org reports through the CEO directly more often than through a CPO sitting between. Each prediction has a specific test you can run in July 2028 to check whether I was right.

### 39 PM AI Agents Deployed: What Stuck, What Died, and Why

Published: 2026-04-27
Canonical: https://falkster.com/blog/39-agents-stuck-vs-died

Forty files in my agent fleet, of which 39 are real agents. Built across 80 days. The fleet is structurally biased toward Decide (7) and Build (7) over Discover (5), which contradicts the popular framing that AI for PMs equals customer interview synthesis. Thirteen agents fire in the 7-to-9 a.m. morning brief slot, which is now the most-contested real estate in the system. Slack is mentioned in 25 of 39 agents, the universal substrate. Thirteen agents are completely orphaned, with zero references from non-agent content anywhere on the site, which is the cleanest "we built it but it didn't stick" signal I have. The single biggest predictor of agent survival is whether a named human owns it. The single biggest predictor of agent death is scope drift, where an agent starts focused and ends as a Frankenstein doing five jobs badly. If I were starting over today I would build seven agents first, all on a single workflow surface, and refuse to add agent number eight until the first seven were fully integrated into a real team's week. I built thirty-nine.

### The Agent Sandbox (PM Version): A Complete User Manual

Published: 2026-04-27
Canonical: https://falkster.com/blog/agent-sandbox-user-manual

The sandbox has nine views, accessed from the top navigation: **Overview**, **PM View**, **Agents**, **Layers**, **Sources**, **Workflows**, **Impact**, **Inbox**, and an **Ask Falkster** chat overlay. The Overview is where the demo starts. The PM View is the most personally useful. The Layers view is where you stop talking about agents as a list and start talking about how the fleet thinks. The Workflows view is the one that changes minds. The Impact view is the one that closes board members.

### The Seven-Agent Reset: A Lean PM AI Fleet, One Agent Per Stage

Published: 2026-04-27
Canonical: https://falkster.com/blog/seven-agent-reset

<StatBlock stats={[

### Auto Bugfix Agent: Zendesk Ticket to Reviewable PR in 11 Minutes

Published: 2026-04-20
Canonical: https://falkster.com/blog/agent-auto-bugfix

The Auto Bugfix Agent watches Zendesk for customer-reported bugs. When one fires, it reproduces the issue, walks the call graph to find the broken code, highlights the specific lines, writes the fix, adds a regression test, and opens a reviewable PR. Engineer on-call spends 5 minutes reviewing instead of 4 hours digging. Signal-to-reviewable-fix drops from days to about 11 minutes.

### Instant Prototype Agent: Customer Request to Prototype in Minutes

Published: 2026-04-20
Canonical: https://falkster.com/blog/agent-instant-prototype

The Instant Prototype Agent listens for customer feature requests across Slack, Zendesk, Gong, and Salesforce. For every validated request, it generates a working prototype in minutes, deploys it to a preview URL, opens a GitHub branch, files a Linear ticket, and drafts a Notion doc with the customer context. The PM goes from "interesting signal" to "clickable artifact the customer can react to" in one cycle, usually the same day.

### Launch Comms Agent: Six Channels Generated on Every Release

Published: 2026-04-20
Canonical: https://falkster.com/blog/agent-launch-comms

The Launch Comms Agent reads a just-shipped feature's full context (PRD, prototype, Linear ticket, release notes, customer story) and generates every piece of go-to-market copy needed in one pass: website hero, LinkedIn post, customer email, in-product banner, public changelog, X thread. Six channels, one voice, under 4 minutes. Product Marketing edits and ships. No more launch days where three channels go out with contradictory copy.

### Signal-to-Ship Cycle Time Agent: Measure PM Velocity Across 7 Stages

Published: 2026-04-20
Canonical: https://falkster.com/blog/agent-signal-to-ship

The Signal-to-Ship Cycle Time Agent tracks how fast every active product change moves through the seven stages of the [PM Operating System](/handbook/product-operating-model): Sense, Discover, Decide, Build, Ship, Measure, Amplify. It runs daily for a snapshot and weekly for a trend digest, names the bottleneck stage, and flags stuck items. Teams that deploy it typically see median cycle time drop 50 to 70 percent in one quarter. It's a meta-agent: it observes the outputs of the rest of your [AI agent fleet](/handbook/ai-agent-army) rather than pulling raw data itself.

### The Ten-Day Dev Loop: Three AI Agents Collapsed Our 8-Week Cycle

Published: 2026-04-20
Canonical: https://falkster.com/blog/ten-day-dev-loop

This is a composite story of one customer request moving through the three new agents in the fleet. Monday afternoon: a CSM at Acme Corp sends a Slack DM. Ten days later: the feature is live, announced across six marketing channels, with the customer's quote on the landing page. The old loop would have taken 8 weeks. The new one took 10 days. Three agents, four clickable artifacts, and a much shorter argument.

### KPI Watchdog Agent: Catch Metric Drops and Ship a Fix Prototype

Published: 2026-04-19
Canonical: https://falkster.com/blog/agent-kpi-watchdog

The KPI Watchdog agent checks your three product KPIs every hour, detects when one moves outside the noise band, investigates the likely cause by clustering it with recent product changes and customer signal, and then ships a clickable prototype of a candidate fix. By the time you read the Slack alert at 8:17 AM, the prototype is already running at a preview URL. Three tiers (Note, Alert, Incident) match alert noise to signal severity. The prototype-builder is the move that separates a 2023 dashboard from a 2026 product-builder agent. Build the v1 in three weekend evenings: one KPI, one collector, one alert. The full version is the highest-ROI agent I've built.

### I Gave My AI Agents a Performance Review. Three Got Fired.

Published: 2026-04-14
Canonical: https://falkster.com/blog/agent-performance-review

If an agent does the work of a team member, it should be managed like one, including being let go when it underperforms. I run a quarterly performance review on every agent in my fleet using four metrics: precision, escalation rate, time-to-output, and trust trend. Last cycle, three agents failed it and I shut them down. Not because they never worked, but because they were confidently wrong in ways I could not catch at review time, which makes an agent worse than useless. The industry talks about agents like magic that either works or does not. The honest version is that agents are workers with track records, and most teams are keeping net-negative ones running out of sunk cost. Here is the scorecard, the coaching loop, and the criteria for when to stop coaching and fire the thing.

### Stakeholder Communication Agent

Published: 2026-04-07
Canonical: https://falkster.com/blog/agent-stakeholder-communication

The Stakeholder Communication agent generates five tailored updates from the same source data: a one-page executive summary, a two-to-three-page board update, an investor update, a team briefing, and a customer email. It runs Friday at 3 PM (team and customer) and monthly (exec, board, investor). Feed it OKR progress, launches, metrics, and risks. The agent applies audience-specific transformations: execs get business impact, boards get risk assessment, investors get unit economics and growth, team gets implementation details, customers get benefit-led announcements. Days of stakeholder writing collapsed into one agent run plus an hour of refinement.

### Retrospective Synthesis and Learning Agent

Published: 2026-04-05
Canonical: https://falkster.com/blog/agent-retrospective-synthesis

The Retrospective Synthesis agent runs every Friday at 5 PM after your retro and turns retro notes into actual learning. It compares today's retro against the last 8 retros, detects recurring themes (the "slow deploys" complaint mentioned 5 times that nobody fixed), tracks action item follow-through, and proposes specific playbook updates. The point is to stop re-discovering the same problems every quarter. Three sprints from now, you'll have a living team playbook that's been updated 6 times with patterns you actually learned. Connect your retro notes repo and run it on this week's session.

### Win/Loss Analysis Agent

Published: 2026-04-05
Canonical: https://falkster.com/blog/agent-win-loss-analysis

The Win/Loss Analysis agent runs bi-weekly Tuesday at 10 AM and reads every won and lost deal from the last two weeks. It pulls from Salesforce (deal size, close reason, sales notes), call transcripts, and win/loss interviews, then extracts seven categories of insight: win patterns, loss patterns, feature insights, segment insights, competitive positioning, pricing insights, and three recommendations. The point is to stop building features nobody asked for and start building the ones that move deals. The output of one run: "We win enterprise on API reliability and lose mid-market on price. Two product fixes would convert 30% of mid-market losses."

### Feature Adoption Tracking Agent

Published: 2026-04-04
Canonical: https://falkster.com/blog/agent-feature-adoption

The Feature Adoption agent tracks adoption curves daily for every feature shipped in the last 3 months. It pulls from Amplitude or Mixpanel, breaks the data into three signals (trial rate, retention rate, cohort breakdown), and flags lagging segments before the trend hardens. When SMB adoption sits at 12% while enterprise is at 40%, you get a root-cause hypothesis and a recommended intervention (in-app tour, email campaign, onboarding tweak). Daily snapshot at 4 PM, deep dive Monday at 9 AM. Connect feature flags so the agent knows when each user got access, then run it on your three most recent launches.

### Automated PRD Generator

Published: 2026-04-04
Canonical: https://falkster.com/blog/agent-prd-generator

The PRD Generator agent runs every Monday at 11 AM, takes the top 2-3 prioritized opportunities, and auto-drafts a full PRD for each one. It grounds requirements in the actual research (interview quotes, journey friction, NPS drivers, support patterns), aligns with the design system, and surfaces technical constraints. The output has eight sections from goals and success metrics through risk and mitigation. The point isn't to skip the PM thinking. It's to skip the blank page and the meetings spent reconstructing context. PRDs take two hours to refine instead of ten hours to write. Connect your opportunity stack, research repo, and Figma, then generate one PRD this week.

### Automated Release Documentation Agent

Published: 2026-04-04
Canonical: https://falkster.com/blog/agent-release-documentation

The Release Documentation agent runs every Wednesday at 5 PM (and on demand for hot-fix releases). It reads the sprint's shipped features from Jira or Linear, the matching PRDs, and your docs repository, then generates four outputs at once: technical release notes, a customer-facing announcement email, a support team briefing with FAQs, and a documentation task list. The point is to stop spending Friday afternoon writing four versions of the same announcement that drift out of sync. Set it up so this Wednesday's release notes write themselves. Edit the customer email tone if needed and publish.

### Tech Debt Impact and Prioritization Agent

Published: 2026-04-04
Canonical: https://falkster.com/blog/agent-tech-debt-analyzer

The Tech Debt Analyzer agent quantifies which debt items are actually slowing velocity. It runs every Monday at 10 AM, analyzes the last 4 sprints of Jira data, correlates each debt area with velocity impact (auth refactor affects 5% of stories vs. legacy data layer affects 40%), and maps upcoming features to the debt they'll touch. The output is a prioritized list scored by (impact × urgency) / effort, with a payback calculation: "Fixing the data layer costs 40 dev-weeks but saves 20 dev-weeks per quarter going forward." Stop guessing which debt matters. Start fixing the data layer first.

### Testable Assumptions Tracker Agent

Published: 2026-04-03
Canonical: https://falkster.com/blog/agent-assumption-tracker

The Assumption Tracker agent converts each prioritized opportunity into three to five testable assumptions and tags each one Not Tested, Testing, Validated, or Invalidated. It runs weekly on Friday at 2 PM, reading the OST, recent interviews, and experiment results. The output is one living document that flags any feature you're about to build on untested assumptions, plus a ranked list of which assumptions are cheapest to test first. I use it to stop teams from shipping four-week builds against assumptions nobody validated. Start by listing five assumptions behind your next sprint commitment.

### OKR Progress and Prediction Agent

Published: 2026-04-03
Canonical: https://falkster.com/blog/agent-okr-tracker

The OKR Tracker agent gives you daily progress scoring and weekly confidence predictions for every key result. Daily at 4 PM, it pulls the latest metrics from Amplitude, Mixpanel, and Salesforce, scores each OKR against the trajectory needed, and flags anything off-track. Friday at 3 PM, it runs the confidence model: "you're at 65% of target with 4 weeks left, current velocity puts you at 82%, here's what would need to change to hit 100%." The point is to stop being surprised at the end of the quarter and catch slippage at week 5 when you can still course-correct. Set up the metrics feeds, then ask which Q2 OKR has the lowest confidence score today.

### Opportunity Prioritization and Synthesis Agent

Published: 2026-04-03
Canonical: https://falkster.com/blog/agent-opportunity-prioritization

The Opportunity Prioritization agent reads the outputs of all five DISCOVER agents (Support Signal Processing, NPS/CSAT Analysis, Interview Synthesis, Journey Mapping, Customer Segmentation) and produces one prioritized opportunity stack every Friday at 11 AM. Three layers of synthesis: cross-dataset validation (opportunities appearing in 2+ reports get high confidence), impact estimation (segment size × WTP × churn reduction), and dependency mapping (fixing A unlocks B). The output is OST-ready: problem statement, evidence, affected segments, estimated impact, rough effort. Going from five scattered reports to one actionable list cuts roadmap planning to a fraction of the time.

### Automated Sprint Planning Agent

Published: 2026-04-03
Canonical: https://falkster.com/blog/agent-sprint-planning

The Sprint Planning agent runs every Monday at 9 AM and converts your prioritized opportunities into a proposed sprint plan in 30 minutes instead of 4 hours. It reads opportunities, tech debt backlog, team capacity, and historical velocity from Jira or Linear, then breaks each opportunity into user stories with story points, fits stories into the sprint respecting capacity and dependencies, and explains every trade-off. The output is ready to load into Jira, with a confidence score on the plan and a list of which stories are most likely to slip. Half a day reclaimed, every sprint.

### Customer Segmentation Agent

Published: 2026-04-02
Canonical: https://falkster.com/blog/agent-customer-segmentation

The Customer Segmentation agent rebuilds your customer segments weekly by clustering on actual behavior, not last year's firmographics. It runs Monday at 9 AM, pulling from Amplitude or Mixpanel for usage, Salesforce for attributes, and the support stack for engagement patterns. The output is 4 to 6 true behavioral segments with size, profile, and shift data, plus flags for emerging segments and at-risk groups. The point is to catch segment evolution while it's still small (the new "AI builders on your API" segment that wasn't there last quarter). Replace your stale company-size buckets with this on next Monday's planning.

### Customer Interview Synthesis Agent

Published: 2026-04-02
Canonical: https://falkster.com/blog/agent-interview-synthesis

The Interview Synthesis agent reads every customer interview transcript from the week (Otter, Fireflies, manual notes) and produces one synthesis report every Wednesday at 10 AM. The report has six sections: recurring themes (3+ interviews), top pains vs. gains, feature mentions, testable hypotheses, contradictions, and segment insights. The point is that signal only emerges across 8+ interviews and manually reading 12 transcripts takes 6 hours. The agent does it in minutes, with customer quotes preserved. Feed in your last 10 transcripts and segment metadata, then ask for the top three testable hypotheses.

### Automated Customer Journey Mapping

Published: 2026-04-02
Canonical: https://falkster.com/blog/agent-journey-mapping

The Journey Mapping agent builds dynamic customer journey maps every two weeks by synthesizing three data sources: session replays (Fullstory or Hotjar), support tickets (Zendesk), and research notes. The output is a map showing actual paths, decision points, drop-off rates, and friction moments, segmented by cohort. The point is to replace static Figma journey maps that go stale with one that updates from real behavior, so "users drop 35% at the data mapping stage" comes with evidence instead of intuition. Bi-weekly Thursday at 9 AM. Start by pulling 50 sessions from one cohort and asking the agent where the friction concentrates.

### Every Type of PM Needs a Different Agent Stack

Published: 2026-04-02
Canonical: https://falkster.com/blog/agents-for-every-pm

Each of the ten common PM roles (Technical, Growth, AI, Product Owner, Platform, Generalist, Product Ops, Data, PMM, E-Commerce) needs a different AI agent stack. The agents aren't one-size-fits-all. A Growth PM lives in Product Health configured for funnel metrics, AARRR, and Red Flag for signup drops. A Platform PM lives in Product Health configured for API uptime and Documentation Gaps for the API docs that ARE their product. A Generalist needs Daily Focus and Weekly Ops Digest. The Product Health agent appears in almost every stack, but configured differently each time. Configuration matters more than capability. Pick three agents that match your actual role, run them reliably, then add one at a time.

### NPS and CSAT Deep Dive Agent

Published: 2026-04-01
Canonical: https://falkster.com/blog/agent-nps-csat-analysis

The NPS/CSAT Deep Dive agent runs a daily snapshot at 4 PM and a weekly deep dive Monday at 8 AM. It pulls survey responses from Typeform, Delighted, or SurveySparrow, maps them to customer segment data in your CRM, and surfaces the actual drivers behind your score. The output isn't "NPS is 47," it's "NPS is 47, enterprise is at 58, SMB is at 31, and the #1 detractor driver is onboarding speed." Daily catches major-customer detractors same day. Weekly shows cohort shifts and feature correlation. Connect your survey platform and ask for the top three detractor themes from last month's responses.

### Automate Support Pattern Detection

Published: 2026-04-01
Canonical: https://falkster.com/blog/agent-support-signal-processing

The Support Signal Processing agent runs daily at 8 AM and pulls every Zendesk or Intercom ticket from the last 24 hours through three parallel analyses: severity clustering by root cause, segment breakdown (Enterprise, Mid-Market, SMB), and trend detection against the 7-day and 30-day baselines. It flags new clusters that didn't exist a week ago, segments with 20%+ ticket volume spikes, and any category that jumped >25% versus baseline. The point is to catch the spike in "API timeout" tickets on day 2 instead of day 6 when a customer escalates. Wire up your support system and the 90-day historical baseline, then run it tomorrow morning.

### Claude Skills Every PM Should Build Today

Published: 2026-03-31
Canonical: https://falkster.com/blog/claude-skills-for-pms

Claude Skills are folders of instructions Claude can reference automatically. They're not Projects (which focus on specific work streams) and they're not custom instructions (which are global). For PMs, the five Skills to build first are: Customer Interview Synthesis, PRD/One-Pager, Stakeholder Update, Competitive Intel, and OKR/Outcome Tracking. Each takes 30 minutes to build and pays back in week one. The triggering hack: add one line to your custom instructions ("check if /prd-skill or /competitive-intel folders apply before responding") so Claude reliably uses them. Start with the PRD skill today. By month's end you'll have your entire PM operating system embedded in Claude.

### Build Your Market Intelligence Agent

Published: 2026-03-27
Canonical: https://falkster.com/blog/agent-market-intel

The Market Intelligence agent runs bi-weekly Wednesdays at 9 AM and aggregates six sources (competitor websites and blogs, Crunchbase, Salesforce win/loss data, Gong sales call transcripts, customer switching signals, analyst reports) into one strategic report. The output is seven sections: competitive landscape changes, feature parity matrix, win/loss patterns, switching signals with ARR at risk, positioning comparison, watchlist, and two-to-three strategic recommendations. The difference from the daily competitive intel agent: this one zooms out to find patterns, not events. Run it for the first time on your top three competitors and see what shifts.

### Agent-to-Agent Dispatch: The Product Org Chart Nobody Is Designing

Published: 2026-03-26
Canonical: https://falkster.com/blog/agent-to-agent-dispatch

The entire market is selling single AI agents to product managers: an agent that writes the PRD, an agent that summarizes research, an agent that drafts the update. That is the easy, demo-friendly part, and it is not where the value is. The value is in the layer nobody is building, where agents hand work to other agents with no human relaying context in between. A listening agent detects an outcome, routes a structured brief to a build agent, a research agent, and a comms agent, and the PM sits above it as the dispatcher who sets policy and owns the quality gates, not the relay who carries messages between tools. The companies that win will not have the most agents. They will have agents that talk to each other. Here is the architecture, and why your org chart is the thing that has to change.

### Build Your Release Checker Agent

Published: 2026-03-23
Canonical: https://falkster.com/blog/agent-release-checker

The Release Checker agent is your final gate before shipping. It runs Thursday at 10 AM (72 hours before a Monday release) and verifies QA test results (Jira or TestRail), documentation status, GTM materials, sign-offs, and feature flag configuration. Each feature gets a status: Go, At Risk, or Blocked. The output is a clear go/no-go recommendation with named blockers, owners, and EOD-Thursday deadlines. The 72-hour buffer is the point. If QA is at 72% Thursday morning, you have Friday to fix it. If you find blockers Monday, you ship with bugs or you scramble. Add this Thursday and your release-day chaos drops to zero.

### Two Weeks of Agent Tuning: What I Learned

Published: 2026-03-23
Canonical: https://falkster.com/blog/playbook-agent-tuning

Day 1 of agent deployment: my morning report had 47 items at 8 AM. 60% noise. By Day 14, the same agents produced 5-7 actionable items I trusted. Six dial adjustments did the work: severity weighting on Day 3, channel filtering on Day 5, customer-tier weighting on Day 7, action-first report structure on Day 10, comparative context on Day 12, and edge-case refinement in week 3-4. The playbook: treat the first two weeks as training, not production. Day 1 output is a first draft. You need to tune before you trust. By Day 14, the report became the first thing I checked each morning. Before that, it was the last.

### Build Your Daily Red Flag Agent

Published: 2026-03-20
Canonical: https://falkster.com/blog/agent-red-flag-detection

The Red Flag Detection agent is the first agent every PM should deploy. It runs daily at 9 AM, scans five sources (Zendesk, PagerDuty, Jira, Slack, Salesforce), and posts one prioritized triage report to a Slack channel using a three-tier model: Critical (must handle today), Warning (monitor, plan response), Info (no action). The report includes recommended next steps with owners and timelines so you go from "what needs my attention?" to action in 10 minutes instead of 60. Over a quarter, it reclaims more than 40 hours of focused time. This is the first agent to wire up. Set it up this week.

### How I Caught a Churn Signal 3 Weeks Early

Published: 2026-03-19
Canonical: https://falkster.com/blog/playbook-churn-signal

Monday morning, my Red Flag agent flagged a tier-1 customer with four support tickets in 48 hours when they normally averaged 1-2 per month. By Friday afternoon, the fourth ticket read "we're evaluating other vendors." The agent did its job: pattern-matched at scale on data I couldn't watch manually. The PM work was the 45-minute call where I caught what the volume actually meant (a deprecation notice they'd missed during an identity provider migration). I built a backward-compatibility shim by Tuesday afternoon, walked through migration on Wednesday morning, and they renewed three weeks later at full terms. Without the agent, that signal lives in the support backlog until CS finally notices renewal is at risk. The agent didn't save the deal. It made sure the PM work happened three weeks early.

### Build Your Product Health Agent

Published: 2026-03-15
Canonical: https://falkster.com/blog/agent-product-health

The Product Health agent runs daily at 4 PM and gives you one synthesized story across seven dashboards: engagement (30% of health score), conversion (25%), retention (20%), performance (15%), and customer satisfaction (10%). It pulls from Amplitude or Mixpanel, Sentry, Zendesk, PagerDuty, and Slack, and produces a narrative summary with green/yellow/red flags. The point is to stop checking seven tools and start the evening knowing whether your product got healthier or sicker today. The agent catches three things you'd miss otherwise: silent performance creep, support volume that correlates with error rates, and underperforming cohorts hidden inside healthy global metrics.

### Setup Guide: Connecting Claude to Your Data Sources

Published: 2026-03-15
Canonical: https://falkster.com/blog/setup-guide-claude

This is my actual production agent setup at Smartcat. Model Context Protocol (MCP) is Anthropic's standard for connecting Claude to external tools. Wire up 8 data sources via MCP servers (Jira, Slack, Zendesk, Salesforce via Databricks, Gong via Weaviate, Google Calendar, Notion, GitHub, Google Drive), schedule agents on cron (7 AM, 9 AM, 4 PM, weekly Monday at 8 AM), and pipe results to Slack via incoming webhooks. The whole setup takes 2 hours. Prerequisites: Claude Pro or Team, Node.js 18+, API keys for your tools, basic terminal comfort. After this, your first agent runs tomorrow morning without you touching anything.

### Setup Guide: Self-Hosted Agents with OpenClaw

Published: 2026-03-15
Canonical: https://falkster.com/blog/setup-guide-openclaw

OpenClaw is the self-hosted agent runtime for teams that can't use Claude's managed setup for compliance, security, or regulatory reasons. Same agents as the Claude setup, same MCP data source connectors, but everything stays inside your firewall on Docker. The architecture has five components: gateway, brain (reasoning engine with pluggable LLMs), memory, skills (MCP servers), monitoring. Prerequisites: Docker, 8GB+ RAM, network isolation, API keys. Run on stable on-prem hardware. Pick OpenClaw if data cannot leave your network or you need custom model routing. Pick Claude's managed setup otherwise. Note: OpenClaw shipped with security vulnerabilities through 2025-2026; audit every plugin and run in a network-isolated environment.

### Build Your Competitive Intelligence Agent

Published: 2026-03-14
Canonical: https://falkster.com/blog/agent-competitive-intel

The Competitive Intelligence agent pulls from seven sources (competitor changelogs, G2 and Capterra reviews, Salesforce deal notes, Gong sales call transcripts, support tickets, Slack channels, industry publications) and ships one synthesized report every other Monday. The report has seven sections, including top five competitive insights, win/loss patterns, and positioning recommendations, all readable in ten minutes. The point isn't to research better than you could manually. It's to systematize the work so nothing falls through the cracks, week after week. I went from 3 hours a week on competitive research to zero. Start with two sources, Gong and Salesforce, and add the rest in week two.

### Build Your Team Triage Agent

Published: 2026-03-11
Canonical: https://falkster.com/blog/agent-team-triage

The Team Triage agent runs twice daily (9 AM and 4 PM) and reads your entire #team-product Slack channel since the last run. It categorizes every message into bugs, feature requests, escalations, and questions, assigns owners by matching against your PM roster, and surfaces threads waiting on response with a "waiting since" timestamp. It also flags ownership gaps (issues nobody claimed), trending topics, and PMs with too many open items. The point is to stop losing issues in a 87-messages-a-day firehose. Setup is 20 minutes. Saves the team 30 minutes a day in triage overhead.

### Build Your Engineering Capacity Agent

Published: 2026-03-09
Canonical: https://falkster.com/blog/agent-engineering-capacity

The Engineering Capacity agent runs Monday at 8 AM and tells you whether the next four weeks of roadmap actually fit in available capacity. It pulls from Jira (assigned work, story points, sprint completion), GitHub (PR throughput), and Google Calendar (PTO and leave), then computes net available capacity per engineer and per team after subtracting PTO, on-call, meetings, and ramp-up. The output flags overloaded engineers, over-capacity teams, understaffed critical projects, and PTO that creates coverage gaps weeks before they hit. The math is explicit: if assigned > capacity, you're behind before the sprint starts. Run it on your current sprint this Monday and check if you're already at 130%.

### Build Your Documentation Gap Agent

Published: 2026-03-06
Canonical: https://falkster.com/blog/agent-documentation-gaps

The Documentation Gap agent catches documentation failures before they ship. It runs daily at 7 AM, scanning three signals: customer confusion (Zendesk tickets, Gong sales call moments), outdated help-center articles, and the release calendar in Jira. The killer output is the release documentation readiness check: a daily report of every Tier 1 feature shipping in the next 14 days alongside the status of its help article. Three features ship next week, one has docs done, two don't. That's not a warning. That's a deadline. Start by connecting Jira and your help center, then ask the agent which features shipped in the last 30 days with no help article. Most teams find a list.

### Build Your Weekly Ops Digest Agent

Published: 2026-03-02
Canonical: https://falkster.com/blog/agent-weekly-ops-digest

The Weekly Ops Digest agent runs Monday at 8 AM and looks at every daily report from the past week to spot patterns your daily agents missed. It calculates week-over-week trends across support tickets, resolution time, escalations, velocity, critical bugs, and deployment frequency. When the same issue shows up in 3+ daily reports, it's a pattern, not a fire. The output is a weekly narrative: what improved, what degraded, what systemic issues to address, and the three priorities for next week. Daily agents catch fires. This one finds the fuel source.

### Build Your Product Ops Agent

Published: 2026-03-01
Canonical: https://falkster.com/blog/agent-product-ops

The Product Ops agent runs daily at 9 AM and synthesizes five data streams (Zendesk tickets, Amplitude usage, Salesforce health, Slack feedback, public reviews) into one operational report ranked by ARR at risk. The output has five sections: top complaints by revenue impact, feature gaps weighted by requesting customers, usage anomalies with red/yellow flags, at-risk customers with reasons and recommended actions, and a revenue-at-risk summary. The point is to stop ranking by ticket volume and start ranking by business impact. The agent catches "Enterprise-X ($500k ARR), usage declining, mentioned competitor in last sales call" on day 1, not day 8 when the cancellation email arrives.

### Build Your Daily Focus Agent

Published: 2026-02-28
Canonical: https://falkster.com/blog/agent-daily-focus

The Daily Focus agent is an AI chief of staff. It runs at 7 AM, reads your Google Calendar, Slack, Jira, and email, and posts the three things that actually matter today to a #daily-focus channel. Each priority is scored on three axes: urgency (0-40), customer impact (0-35), executive visibility (0-25). The top three become your day. The report also surfaces invisible blockers, executive commitments coming due, and Slack threads waiting on your decision. Setup is 30 minutes. The point is to stop checking five tools every morning and start the day knowing what to focus on.

### Build Your Customer Commitment Agent

Published: 2026-02-25
Canonical: https://falkster.com/blog/agent-customer-commitments

The Customer Commitment agent tracks every promise sales, customer success, and leadership make to customers, cross-references each one against the actual roadmap, and tells you which commitments are overdue, at risk, or were never planned. It runs twice a week, Tuesday and Friday at 9 AM, reading Gong call transcripts, Salesforce deal notes, CS meeting notes, and Slack. The output is a four-category report (overdue, at-risk, no roadmap item, new this week) with ARR exposure attached to each. I built this after losing a $500k deal to a forgotten promise. The quarterly commitment accuracy score (% delivered on time) is the most uncomfortable number it surfaces.

### Put Stakeholder Updates on Autopilot: The Complete Setup Guide

Published: 2026-02-22
Canonical: https://falkster.com/blog/stakeholder-update-autopilot

The stakeholder update is one of the highest-friction, lowest-value tasks in a PM's week. Build an agent that runs every Monday morning, pulls data from six systems (GitHub, Jira, analytics, Zendesk, Salesforce, Slack), and generates three audience-tailored updates: executive summary (250 words, data + decisions needed), team update (400 words, what shipped + what's next), board update (300 words, MRR + retention + risks). Three setup options included: manual Claude in browser (5 minutes), Zapier with Claude API (30 minutes), or specialized AI tool. Saves 5-6 hours weekly. Stakeholders get better updates because you're synthesizing real data instead of writing from memory. Start with manual review then automate auto-send after 3-4 weeks of trust. The one-page format the autopilot should draft into is [the exec update template](/blog/exec-update-template).

### Build Your PM Issues Agent

Published: 2026-02-20
Canonical: https://falkster.com/blog/agent-pm-issues

The Daily PM Issues agent runs every weekday at 7 AM, before your standup, and surfaces three things: features at risk of missing their ETA, customer commitments coming due this week, and cross-team dependency failures. It pulls from Jira, Salesforce, three Slack channels (#team-product, #engineering, #customer-success), and PM meeting notes. The report has five sections including a one-paragraph executive summary: the single most urgent issue and what to do about it in the next four hours. I built this after watching a $200k deal walk because an engineering blocker sat in Slack for two weeks and never reached the PM who owned the customer. Run it before next Monday's standup.

### Build Your Product Health Dashboard Agent

Published: 2026-02-19
Canonical: https://falkster.com/blog/agent-product-dashboard

The Product Health Dashboard agent runs every Tuesday at 9 AM and goes deeper than daily metrics. It pulls 3 weeks of data from Amplitude or Mixpanel and analyzes four dimensions: feature adoption curves (day 1, 3, 7, 14, 30 trajectories vs. baseline), cohort retention (this week vs. 4 and 8 weeks ago), engagement depth (surface vs. regular vs. power users), and performance trends from DataDog or Sentry. The point is to distinguish "not valuable" from "not discoverable." A feature with 20% adoption that's flat could be a discovery problem with power-user retention 40% higher than overall. The dashboard tells you which. This is where quarterly strategy gets shaped.

### Build a Customer Feedback Pipeline in One Afternoon

Published: 2026-02-18
Canonical: https://falkster.com/blog/customer-feedback-pipeline

This is a one-afternoon setup that aggregates customer feedback from Zendesk, Slack, and app store reviews into a single AI-processed digest. The architecture is three steps: extract (Zapier or a Python script hitting Zendesk and Slack APIs), process (Claude's API with a structured prompt that returns category, sentiment, urgency, segment, quote, and action item as JSON), digest (post to Slack at 7 AM or email). Total setup time is 2-3 hours, mostly waiting on API keys. Ongoing time is 5 minutes a day. By Day 2 morning, you're seeing customer patterns before they become problems. Code snippets are inline. Pick four sources where your customers actually talk to you and ship it tonight.

### Build Your Release Readiness Agent

Published: 2026-02-15
Canonical: https://falkster.com/blog/agent-release-readiness

The Release Readiness agent runs every Wednesday at 8 AM and audits five critical inputs: PRD approval status, release notes registration, ETA accuracy, feature flag coverage, and GTM readiness. The output is an 11-section report including a critical-gaps dashboard, a full release inventory, PM action items with named owners, and a go/no-go recommendation. The agent catches the three things that derail every launch: missing PRDs past deadline, ETAs that silently slip past release date, and features that ship without a release notes entry. By week three of running this, launches are boring. Set up the source MCP connections (Jira, GitHub, Notion) and run it on next Wednesday's release slate.

### Build Your GTM Release Monitoring Agent

Published: 2026-02-12
Canonical: https://falkster.com/blog/agent-gtm-monitoring

The GTM Release Monitoring agent runs daily at 7 AM and tells you which features are actually ready to ship across five readiness gates: discovery compliance, beta program standards, GTM materials, beta feedback timing, and compliance signoff. The report uses a launch countdown view (green/yellow/red) per upcoming launch. The point is to stop tracking "is the code done" and start tracking "can we sell it, support it, succeed with it." Connect Jira, Slack, and Google Drive. Run it on the three features closest to launch and see which one is actually ready.

### Set Up a Competitive Intelligence Agent in 30 Minutes

Published: 2026-02-09
Canonical: https://falkster.com/blog/competitive-intel-agent

A competitive intelligence agent is an AI agent that watches your top 3-5 competitors across pricing, features, messaging, hiring, and reviews, and delivers a structured weekly Slack brief every Friday. Setup is 30 minutes: list competitors, pick signals, paste a weekly-aggregation prompt into Claude or ChatGPT, schedule it via Zapier or Make. Cost: roughly $20/month. Output: a one-page brief that highlights the three moves that matter and ignores the rest. For the full agent fleet this fits into, see [Your AI Agent Fleet](/handbook/ai-agent-army). For the full template prompts in copy-paste form, grab the [downloadable artifact](/artifacts/competitive-intel-prompts.md) at the bottom of this post.

### Build Your Roadmap Progress Agent

Published: 2026-02-08
Canonical: https://falkster.com/blog/agent-roadmap-tracker

The Roadmap Progress Tracker agent runs every weekday at 9 AM and cross-references your roadmap against engineering reality. It checks four systems: Jira or Linear (roadmap status), engineering tickets (last activity), GitHub commits (real code activity), and PM updates in Slack. The output is a six-section report with discrepancy alerts ("in progress" with zero commits in 5 days), stale tickets, missing engineering tickets, ETA accuracy trends, ownership gaps, and a daily summary. The point is to stop operating on faith and start operating on data. Roadmaps and engineering work live in separate systems and drift weekly. Run it tomorrow and check how many of your "in progress" items had a commit this week.

### Build Your Executive Report Agent

Published: 2026-02-06
Canonical: https://falkster.com/blog/agent-executive-report

The Executive Report agent writes your weekly Monday morning leadership brief automatically. It runs at 7 AM, aggregates data from Salesforce, Jira, Amplitude or Mixpanel, and Zendesk, and produces a structured 3-minute read: executive summary, top 3 wins, top 3 risks ranked by business impact, roadmap health, metrics snapshot, customer escalations, resource concerns, and the specific decisions needed from leadership. The shift it creates is that your CEO walks into the sync already prepared, so the meeting is about decisions, not status. Start by listing the three risks you'd flag this Monday if you had a clean briefing template.

### Your 'AI Agent' Is Probably Just a Cron Job

Published: 2024-12-31
Canonical: https://falkster.com/blog/agents-vs-workflows-vs-automations

Most things marketed as AI agents are actually automations or workflows with a chatbot bolted on. The distinction matters because it changes your architecture, pricing, reliability story, and customer promise. Automations follow rules. Workflows add a language model at fixed steps in a deterministic pipeline. Agents decide what to do next based on context, adapt to unexpected input, and produce variable outputs. Most enterprise AI value today lives in the workflow tier. Agents get the headlines, but well-built AI workflows get the work done. Don't oversell a cron job as an agent because the word sounds better in the pitch deck.

### Many AI Agents Are Actually Workflows or Automations in Disguise

Published: 2024-12-31
Canonical: https://falkster.com/blog/medium-agents-vs-workflows

Many vendor "AI agents" are actually automations or workflows wearing an expensive suit. The distinction matters because each behaves differently. Automations are rule-based and deterministic. AI workflows add language model capabilities at fixed steps in a deterministic process. Real agents are autonomous, adaptive, and non-deterministic, with continuous learning. Five characteristics separate real agents: autonomy, adaptability, contextual understanding, skill composition, continuous learning. Most enterprise AI value today lives in the workflow tier (Wave 2, 2024-2026). Agentic AI (Wave 3) is emerging. Know which wave you need before believing the vendor hype.

### AI Agents and the Future of Work: A Pixar-Inspired Journey

Published: 2024-10-26
Canonical: https://falkster.com/blog/medium-pixar-ai-agents

A narrative-style exploration of what happens when AI agents become coworkers. Casey watches her mid-sized SaaS company deploy an AI agent platform, watches the agents flounder for two months because they're "starving" without a knowledge graph, then watches the team rebuild the workflow around humans plus agents. By month six the metrics are extreme: 40% faster project delivery, 70% less decision latency, 85% less rework, employee satisfaction up. The lesson lands somewhere else. Companies need to invest in their knowledge graph, stay obsessed with outcomes over outputs, retrain people instead of cutting them loose, and choose ethics over expedience. This is already happening at the companies paying attention.
