Should you ship an AI feature that hallucinates 3% of the time?

THE SHORT ANSWER

Not on the aggregate. First split the eval by traffic slice, because the 3% concentrates: a feature of mine scored well on average across 200 real cases and failed badly on one slice that was about a fifth of actual traffic, which a failure-mode breakdown would have called a small share of unsupported claims. Then write down who measures it in week six, because a sentiment classifier I shipped went from 89% on held-out data to 61% in production within six weeks. Marily Nika's consequence, detectability, and recovery framework holds; the change is that detectability is a product decision, a named eval set with a threshold, a dashboard before launch, and a kill condition, not a fact about whether users can spot the error.

Marily Nika's interview question is the right one, and her framework survives a real launch. Which Slice Holds the 3%? adds the two numbers that change what a candidate should answer: the slice where the 3% concentrates, from The Eval That Caught What the Demo Did Not, and the week-six drift from 10 AI Agents I Built That Failed. The Honest Retrospective., and moves detectability from the user's side of the table to the product's, using the instrumentation contract in Ship With Observability or Don't Ship.

It belongs to the AI Product Management argument: the parts of the job nobody trained for are the ones that decide whether a number is safe to ship on.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-09-29 · 1 min read