A wrong answer from a retrieval-backed assistant has three possible causes, and an accuracy number doesn't separate them. Search didn't return the fact. The model had the fact and lost it. Or search returned the fact along with an older version, and the model picked the wrong one.
The oracle run separates the first two. It comes from Lesson 03 of the Becoming Forward Deployed Engineer course on Substack, which builds a fake model that only repeats the expected fact when it appears in the context. On the course's seeded claims corpus the oracle scores 60% and a real local model 55%, so most of the error sits in search and a better model could recover about five points.
The third cause is mine to add, from an internal docs agent that answered from outdated wiki pages and produced two HR incidents. The oracle would have scored those retrievals as hits. Counting the cases where a conflicting document exists is manual work, and it's the number that tells you whether the next fix is code or content.
The full method, with the five steps and where the result goes in a forward deployed engagement, is in Run the Oracle Before You Swap the Model. The eval loop it plugs into is in Building Evals: Error Analysis, Eval Types, the Loop, the Gate.