ExecutionNew·Falk Gottlob··7 min read

Which Slice Holds the 3%?

Marily Nika asks what is inside a 3% hallucination rate. Mine sat in one slice, a fifth of traffic, and an 89% was 61% by week six. Ask which slice, and when.

hallucinationAI evalsship decisionMarily NikaAI Product Academyslice evalsproduction driftkill conditionobservabilitythe rewrite
Helpful?

Execution Falkster cover on teal blue: a cream pie chart with one wedge pulled out in green under a dark magnifying glass, beside a flip calendar with three torn page corners hanging below it.

Marily Nika published an interview question last week that I would put in front of any AI PM candidate. Your AI feature hallucinates 3% of the time and users love it. Do you ship? Her answer is the right shape. I want to add the two numbers that made me stop trusting the shape, and move one of her axes from the user's side of the table to ours.

The short version

Marily Nika's flashcard evaluates a 3% hallucination rate on consequence, detectability, and recovery, and breaks the 3% into failure modes with a launch bar per mode. Two numbers from my own launches change the question. First, the 3% lives in a slice: a feature of mine scored well on average across 200 real cases and failed badly on one slice that was a fifth of actual traffic, which a failure-mode breakdown would have called a small share of unsupported claims. Second, the 3% has a date: a sentiment classifier went from 89% on held-out data to 61% in production within six weeks, before any user had time to stop checking. And detectability is not a fact about the user, it is a product decision: a named eval set, a pass threshold, a dashboard before launch, and a kill condition, running on production traffic by slice. The rewrite: not "what is inside the 3%," but which slice holds it, and who measures it in week six.

What the flashcard says

Her point is that 3% by itself tells you almost nothing. Three percent of restaurant recommendations naming a place that does not exist is a different product from three percent of medication answers inventing a dose. A 97% product whose failures are obvious and harmless can be safer than a 99.5% product whose failures are subtle, confident, and impossible for the user to see.

So she evaluates on three dimensions. Consequence: is the failure harmless, annoying, costly, or dangerous, and can it propagate into another system. Detectability: can the user catch it, are there sources, is there a confidence cue, is the user asking precisely because they cannot verify. Recovery: retry, edit, undo, escalate, revert to a deterministic path.

Then she breaks the 3% apart. Her example split is 1.4% outdated facts, 0.8% unsupported claims, 0.5% fabricated citations, 0.2% entity confusion, 0.1% high severity, and the launch bar becomes a threshold per mode rather than one accuracy number. She asks what users do with the answer downstream, the blast radius. And she notes that trust changes behavior: early users scrutinize every answer, and six months later they may assume the system is correct.

All of that survives contact with a real launch. What follows is what I would add from mine.

The first number: the 3% sits in a slice

The Eval That Caught What the Demo Did Not is the launch that taught me this. We built an AI feature, demoed it on three hand-picked cases, and the room wanted to ship. I ran the eval I had written before building: the same feature against 200 real cases, split into named slices by the kind of customer and the kind of request.

On average it scored well. On one slice, roughly a fifth of our actual traffic, it failed badly, in a way that would have reached customers.

Here is the part that matters for Marily's framework. By failure mode, that slice's problems were mostly unsupported claims, a modest share of the total. Nothing in a failure-mode breakdown would have flagged it. By slice, it was one in five customers getting the broken version every time. We did not ship. We fixed the slice, re-ran, and shipped a week later on a number instead of a feeling.

The 3% is never spread evenly across your users. It concentrates. The consequence question, the detectability question, and the recovery question all have different answers on the slice where the failures live, and you cannot ask them until you know which slice that is.

The second number: the 3% has a date

Marily's trust point is right, and it arrives later than the other thing that moves.

In 10 AI Agents I Built That Failed. The Honest Retrospective., agent number two is a customer sentiment classifier that scored 89% on held-out data and 61% in production within six weeks. The held-out set was English-heavy. Customers writing in other languages were being scored as angry regardless of what they wrote. The eval did not cover the production distribution.

Six weeks. Nobody had stopped checking yet. The number moved because the traffic was not the test set, and it would have moved on a feature with a 0.5% launch rate just as surely as on one with 3%.

The fix was language-segment evals every week, and a rule I have kept since: never trust an aggregate accuracy number without the per-segment breakdown. Which is the same rule as the first number, seen from the other end. Segment at launch to find where the 3% is. Segment every week to find out what it has become.

The axis to move: detectability is ours

Marily's detectability dimension asks whether the user can catch the error. Can they verify it independently, do we show sources, do they know what good looks like. Those are real questions and the answers should raise or lower the bar.

But whether the user can catch it is a fact about the user. Whether the product catches it is a decision, and the decision is the PM's.

The handbook chapter Ship With Observability or Don't Ship is where I keep that decision. No feature leaves staging without a one-page instrumentation contract written before any code: the success metric, the leading indicators, a cost meter per successful action, a named eval set with a pass threshold, the trace points, a dashboard URL that exists before launch even if it is empty, and a kill condition, the metric, the threshold, and the period after which the feature is reviewed for deprecation. If any of the seven is missing, the feature waits.

That contract is what turns detectability from a hope about users into a number on a dashboard. The eval that found the slice runs against production traffic, by slice, every week. When the week-six number moves, it moves on a chart with a threshold under it, and the kill condition says what happens next. The Receipt Gate is the enforcement half of the same idea: nothing in that pipeline asks the model whether it did the work, because the point of measuring is that nobody has to trust anybody.

The rewrite

Marily's question is "what is inside the 3%." Mine is two questions that come before it.

Which slice holds the 3%? Split the eval by the slices that matter to revenue and to harm, and find the one where the failures live. The launch bar is per slice, and the mitigations she lists, narrow the scope, require citations, add a confidence threshold, route to a human, limit autonomy, are applied there, not everywhere.

Who measures it in week six? Name the eval set, the threshold, the dashboard, and the kill condition before the demo. If the answer is "the users will tell us," you have written detectability as a property of your customers, and they did not sign up for that job.

What to do this week

Take the AI feature closest to shipping. Split its eval set by the three slices that matter most to revenue, and run it. If one slice is not on the dashboard yet, you do not have a 3%. You have an average, and the average is where the demo lives.

This one is part of the AI Product Management argument: AI collapsed the cost of the parts PMs were trained to be good at, and left the parts nobody trained for. Deciding which slice a number hides in, and writing the kill condition before the launch, is one of those parts.

Related answer: Should you ship an AI feature that hallucinates 3% of the time?

Sources: AI PM Flashcard #2: When Is It Safe to Ship Hallucinations?, Marily Nika, AI Product Academy, September 22, 2026. The Eval That Caught What the Demo Did Not, falkster.com, August 12, 2026. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026.

Share this post

Frequently asked

Should you ship an AI feature that hallucinates 3% of the time?+

Not on the 3%. First find which slice of traffic holds it. In one launch I ran, a feature scored well on average across 200 real cases and failed badly on one slice that was about a fifth of actual traffic. Then decide who measures it in week six, because a classifier of mine went from 89% on held-out data to 61% in production in six weeks. Ship on a launch bar per slice plus a kill condition, written before the demo.

What is Marily Nika's framework for shipping hallucinations?+

Consequence, detectability, and recovery. Her post breaks the 3% into failure modes (her example: 1.4% outdated facts, 0.8% unsupported claims, 0.5% fabricated citations, 0.2% entity confusion, 0.1% high severity), asks what users do with the answer downstream, and sets the launch bar per failure mode instead of one accuracy threshold. It is a good framework. The two things I add are the slice and the week.

Why break a hallucination rate down by traffic slice and not only by failure mode?+

Because the failures cluster by who is asking, not only by what went wrong. A breakdown by failure mode called my broken slice a small share of unsupported claims. A breakdown by slice showed one in five customers getting the bad version. The fix was on the slice, and a failure-mode view would have shipped it.

How fast does an AI feature's accuracy drift after launch?+

Faster than the trust curve. My sentiment classifier scored 89% on an English-heavy held-out set and 61% in production within six weeks, because non-English customers were being scored as angry regardless of content. Marily Nika's post says users may stop checking after six months. The number moved in six weeks. The fix was language-segment evals every week and never trusting an aggregate without the per-segment breakdown.

What should be written down before an AI feature ships?+

The instrumentation contract from the handbook: one success metric, leading indicators, a cost meter per successful action, a named eval set with a pass threshold, trace points, a dashboard URL that exists before launch even if empty, and a kill condition stating the metric, the threshold, and the period after which the feature is reviewed for deprecation. If any of the seven is missing, it does not leave staging.

THE SHORT ANSWER

PART OF

AI Product Management

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.