How do you tell whether a wrong AI answer is a retrieval problem or a model problem?

THE SHORT ANSWER

Run the same golden set twice. First with an oracle, a stand-in for the model that cannot reason: it checks whether the expected fact is in the context search returned, repeats it if so, and refuses otherwise. Its score is the ceiling retrieval imposes, so 100% minus the oracle's score is retrieval's share of the errors. Then run the real model. The gap between the oracle and the real model is the model's share. In Lesson 03 of the Becoming Forward Deployed Engineer course, on a seeded teaching corpus with 42 hand-picked cases, the oracle scored 60% and a local model 55%: about 40 points lost to search and about 5 to the model. The oracle has one blind spot. If search returns the right fact together with an outdated version of it, the oracle counts a hit. My internal docs agent failed exactly that way, weighting a 2021 page and a 2024 page equally. So add a third number, the version share: of the cases the oracle gets right, how many have a document in the corpus that answers the same question differently.

A wrong answer from a retrieval-backed assistant has three possible causes, and an accuracy number doesn't separate them. Search didn't return the fact. The model had the fact and lost it. Or search returned the fact along with an older version, and the model picked the wrong one.

The oracle run separates the first two. It comes from Lesson 03 of the Becoming Forward Deployed Engineer course on Substack, which builds a fake model that only repeats the expected fact when it appears in the context. On the course's seeded claims corpus the oracle scores 60% and a real local model 55%, so most of the error sits in search and a better model could recover about five points.

The third cause is mine to add, from an internal docs agent that answered from outdated wiki pages and produced two HR incidents. The oracle would have scored those retrievals as hits. Counting the cases where a conflicting document exists is manual work, and it's the number that tells you whether the next fix is code or content.

The full method, with the five steps and where the result goes in a forward deployed engagement, is in Run the Oracle Before You Swap the Model. The eval loop it plugs into is in Building Evals: Error Analysis, Eval Types, the Loop, the Gate.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-09 · 2 min read