agents·Falk Gottlob··updated ·4 min read

Testable Assumptions Tracker Agent

Convert opportunities into testable assumptions. Track validation status weekly. Know which assumptions are holding up your roadmap.

agentsdiscoveryexperimentation
Helpful?

Pastel title-card cover for the falkster.com post: Testable Assumptions Tracker Agent.
Try it live
See this agent running in the sandbox

Stream a simulated run, inspect the notifications it would send on Slack and email, and see exactly where it sits in the 7-stage PM OS flow. No password required.

The short version

The Assumption Tracker agent converts each prioritized opportunity into three to five testable assumptions and tags each one Not Tested, Testing, Validated, or Invalidated. It runs weekly on Friday at 2 PM, reading the OST, recent interviews, and experiment results. The output is one living document that flags any feature you're about to build on untested assumptions, plus a ranked list of which assumptions are cheapest to test first. I use it to stop teams from shipping four-week builds against assumptions nobody validated. Start by listing five assumptions behind your next sprint commitment.

You're about to build a feature that takes 4 weeks. You haven't actually tested whether customers want it. You're assuming.

The hard part is managing assumptions across a team. Someone tested customer pain in interviews, but you built the feature assuming a different solution. Someone ran a Slack survey that said "add feature X," and you never validated that users would actually use feature X at the volume required to move the needle.

The Assumption Tracker agent converts opportunities into testable assumptions, tracks validation status, and flags which assumptions are still holding up your roadmap. Every Friday it updates a living document. Here's what we're assuming, here's what we've tested, here's what's still uncertain.

How it works

The agent works with two inputs: opportunities and experiment results.

It takes a prioritized opportunity like "SMB customers struggle with slow onboarding" and breaks it into testable assumptions:

  • Assumption 1: "SMB customers want faster onboarding" (validated: interviews, support data)
  • Assumption 2: "They'd use an automated setup flow if available" (not tested)
  • Assumption 3: "A 30-minute setup vs. 2-hour setup would reduce churn by 5%+" (not tested)
  • Assumption 4: "Automated setup would take engineering less than 3 weeks" (not tested)

Each assumption carries a status: Not Tested, Testing, Validated, Invalidated. The agent tracks when it was tested, what the result was, what changed, and what the next test should be.

And it flags risk. If you're building a 4-week feature on assumptions nobody has tested, it says so. "You're assuming users will adopt this, but you haven't tested the adoption."

Data sources and setup

Complete the Claude setup guide first. You'll need:

  • OST / Prioritized opportunities: From the opportunity prioritization agent
  • Research documents: Interviews, surveys, user testing results
  • Experiment tracking: Mixpanel experiments, feature flags, A/B tests, or Notion database
  • Engineering estimates: Feature specs and effort estimates

Schedule: weekly Friday at 2 PM. It also updates when new experiment results come in.

The Claude Prompt

You are tracking assumptions and their validation status.

Here are our prioritized opportunities:
[OPPORTUNITIES LIST]

Here's our research and validation history:
[RESEARCH DATA: interviews, surveys, test results from past 3 months]

Here are current and recent experiments:
[EXPERIMENT DATA: results, holdout groups, metrics, conclusions]

Please analyze and report:

1. **For each top-10 opportunity, break down:**
   - The core problem statement
   - 3-5 testable assumptions (specific and measurable, not vague)
   - For each assumption, current status: Not Tested / Testing / Validated / Invalidated
   - Evidence for each status (what data supports it?)

2. **Assumption Validation Ranking**
   - Which untested assumptions would be quickest to validate?
   - Which assumptions carry the most risk if wrong?
   - Which assumptions matter most for your next sprint?

3. **Experiment Recommendations**
   - For each untested assumption, suggest how to test it
   - Rough estimate: days to run the test? sample size needed?
   - What's the minimum viable validation?

4. **Red Flags** (IMPORTANT)
   - Are you about to build a feature based on untested assumptions?
   - Which assumptions, if invalidated, would kill the whole opportunity?
   - Are there conflicting assumptions between different opportunities?

5. **Assumption Evolution**
   - Which past assumptions were validated? Which were invalidated?
   - Did you learn anything that should change your roadmap?

Format as a tracker: opportunity → assumptions → status → evidence → next test.

What you get

You stop building blind. You know which features rest on tested assumptions and which rest on speculation. You get red flags on the high-risk ones before you start building. You know which assumptions to test first and how, instead of guessing. And you keep a memory of what worked and what didn't, so you don't repeat the same mistake next quarter.

The outcomes I've seen: weeks saved by testing before building instead of after, better adoption because features are validated before launch, and a team that talks about risk and evidence in the same language.


For the full agent fleet and scheduling details, see Your AI Agent Fleet.

Share this post

Also on Medium

Full archive →

Frequently asked

What does the Assumption Tracker Agent do every Friday?+

It reads your prioritized opportunities, recent interview notes, and experiment results, then updates a living document showing every assumption behind your top ten opportunities. Each assumption gets a status: Not Tested, Testing, Validated, or Invalidated, plus the evidence behind that status and a recommended next test.

How does the agent convert an opportunity into testable assumptions?+

It breaks a prioritized opportunity into three to five specific, measurable assumptions. For 'SMB customers struggle with slow onboarding,' the agent produces assumptions like 'SMB customers want faster onboarding' (validated by interviews), 'They would use an automated setup flow' (not tested), and 'A 30-minute setup would reduce churn by 5% or more' (not tested). Each assumption gets its own evidence trail.

What happens when you're about to build on untested assumptions?+

The agent flags it. If you're scoping a four-week feature and three of the five underlying assumptions have never been tested, the report surfaces a red flag with a list of the untested assumptions and what the minimum viable test would be for each. This stops teams from shipping against assumptions nobody validated.

How does the agent rank which assumptions to test first?+

Two dimensions: how quickly can you test it (days to run, sample size needed, minimum viable test) and how much risk does it carry if wrong (would this assumption being wrong kill the whole opportunity?). The intersection of 'fast to test' and 'high risk if wrong' is where the agent tells you to start.

What inputs does the Assumption Tracker Agent need?+

Three: your prioritized opportunities from the opportunity prioritization agent or OST, your research and validation history (interview notes, surveys, test results from the past three months), and your current and recent experiment data (A/B test results, feature flag holdouts, conclusions). It runs weekly on Friday at 2 PM.

PART OF

Enterprise AI Agents

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.