
The answers are wrong 45% of the time, and the room is debating which model to try next. Nobody in that meeting knows what the 45% is made of.
A new Substack course called Becoming Forward Deployed Engineer has the cheapest way out of that meeting I've read. The author is unnamed. Lesson 03 is the one to steal.
The short version
When a retrieval-backed assistant gives a wrong answer, either search failed to hand over the fact or the model had the fact and lost it. Lesson 03 of Becoming Forward Deployed Engineer separates the two with an oracle: a fake model that cannot reason, only repeat the expected fact if it is in the context. Its score is the retrieval ceiling. In the course's seeded claims example the oracle scores 60% and a real local model 55%, so about 40 of the 45 lost points are search and about 5 are the model. I add one column the oracle cannot see, from an agent of mine that failed: the fact was retrieved, and so was its outdated twin. Three shares on one line, retrieval, model, and version, before anyone proposes a bigger model.
What the lesson builds
The course runs a teaching engagement at an invented insurer with a seeded corpus, so every number below is a demonstration number. That's fine. The method is what transfers.
Lesson 1 measures the customer's current keyword search before any AI goes near it: 57% of questions answered in the top five results. Then it splits the average. Lookups by claim number, 100%. Plain-language questions, 15%. Its own summary line is the one I'd put on a wall: the average hid the finding.
Lesson 03 adds a model on top and builds the eval before the feature. Four pieces.
A golden set of 42 cases, picked by hand. The first version sampled 40 at random and, by luck, none landed on a claim whose reserve had never been set, roughly one claim in seven. So the suite couldn't tell a system that says "not yet established" from one that invents a number. Two cases went in on purpose, one of them a claim that doesn't exist, where the only correct answer is a refusal.
Two scorers that check for the fact and ignore the wording. Compare against a model-written reference answer and you're grading style. Reword the prompt and the score moves.
The oracle. It "cannot reason, read or guess." It looks for the expected fact in the prompt and repeats it, or refuses.
And a gate that fails the build if any of four metrics drops more than two points, run against the stand-in so it works offline in about a second.
Then the two runs, side by side:
| Oracle | llama3.1:8b | |
|---|---|---|
| Answer in context (answerable cases) | 59% | 59% |
| Answered correctly | 60% | 55% |
| With a claim number | 100% | 95% |
| Plain language | 15% | 10% |
| Lost after retrieval | 0.0% | 4.9% |
A perfect model scores 60%. The lesson's line: 40 points are gone before the model is asked anything. A frontier model buys back at most a handful, and costs the conversation with the CISO about claim data leaving the network.
The lesson says the whole thing takes about two hours. I believe it.
The column the oracle cannot see
Here's what I'd add, and it comes from a failure I've already published.
Number 5 in 10 AI Agents I Built That Failed. The Honest Retrospective. is an internal docs agent. Classic retrieval setup, answers with citations. Retrieval worked. It found the page. It also found the 2021 version of the page, gave the two equal weight, and answered confidently from the wrong one. Two HR incidents traced back to it. What I wrote down afterward was that source quality matters more than model quality, and that I should have spent 80% of the build on curation and versioning.
Run the oracle against that agent and it scores a hit. The expected fact is in the context. The oracle's whole job is to say so.
So "answer in context" has two populations inside it. The fact arrived alone, or it arrived with a contradiction. A real model facing the second kind will sometimes pick wrong, and the miss lands in "lost after retrieval," which the sheet has labeled the model's share. Someone will read that row and go shopping for a model. Wrong fix again, one column over.
The seeded corpus in the course doesn't have this problem, since its facts are unique by construction. A customer's SharePoint does.
The drop
One sheet, three shares. An afternoon in week one.
- Pick about 40 questions by hand, with the people who ask them. Their answers, their sign-off. Include a question with no answer anywhere in the corpus.
- Run the oracle. 100% minus its score is retrieval's share.
- Run the real model on the same cases. Oracle minus model is the model's share.
- For every case the oracle gets right, count the documents that answer the same question differently. Old policy versions, a superseded price list, the wiki page nobody archived. Cases with one or more are the version share. This step is manual and it's the one that hurts.
- Write the three shares on one line. Retrieval, model, version. That line goes to the meeting. The model debate can start after it.
Then commit the oracle run as the baseline and let the gate block on it. Measure the real model when you choose to, not on every build.
A caution on step 4. I don't have a benchmark for what version share is normal, and I wouldn't trust one. Count your own.
Where it goes after week one
For a forward deployed engineer this is the first artifact that outlives the engagement. The packet in Absorb Pain, Excrete Product: The FDE Handback carries the customer's eval rows back to product. Put the three shares at the top of that page. Retrieval share is a product problem and probably generalizes. Version share is the customer's content problem, and somebody at the customer has to own it by name, or it comes back in month two looking like a regression.
It's also step three of Building Evals: Error Analysis, Eval Types, the Loop, the Gate, fixed inputs, baseline, one change, with the baseline chosen so it tells you which one change to make.
This is a standalone piece next to The FDE Transition, and it belongs to the argument on AI Product Management: the build got cheap, and what's left is knowing which part to build. An accuracy number tells you you're unhappy. It doesn't tell you where.
This week: take the last "we need a better model" request on your team and ask for the oracle run. If nobody can produce one, that's the two hours.
Related answer: How do you tell whether a wrong AI answer is a retrieval problem or a model problem?
Sources: Becoming Forward Deployed Engineer (author unnamed), "Lesson 03: Your First Eval Before Your First Feature," October 6, 2026, and "Lesson 1: A Vague Ask Is Not a Spec Until You Measure It," October 1, 2026, both read in full; the figures are from the course's seeded teaching corpus. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026, for the internal docs agent.
Frequently asked
How do you tell whether a wrong AI answer is a retrieval problem or a model problem?+
Run the eval twice. First with an oracle, a stand-in for the model that cannot reason: it checks whether the expected fact is present in the context search returned, repeats it if so, and refuses otherwise. Its score is the ceiling retrieval imposes. Then run the real model on the same cases. The gap between the ceiling and 100% is retrieval's share of the errors, and the gap between the real model and the ceiling is the model's share.
What is an oracle in an AI eval?+
A deterministic stand-in for the model that is perfect at reading and does nothing else. In Lesson 03 of the Becoming Forward Deployed Engineer course it parses the question, looks up the expected fact, and returns it only if that fact appears in the prompt. Because it never reasons or guesses, its accuracy is the best any model could reach given the documents retrieval handed over.
What did the oracle show in the course's example?+
On a seeded teaching corpus of insurance claims with 42 hand-picked cases, the oracle answered 60% correctly, with the answer present in context for 59% of answerable cases. A local llama3.1:8b model answered 55%. The lesson's reading: of 45 points of error, about 5 belong to the model and about 40 belong to search. These are the course's demonstration numbers, not a customer's.
What can an oracle eval not see?+
A fact that is in the context alongside an older version of itself. My internal docs agent retrieved pages from 2021 and 2024 and gave them equal weight, and two HR incidents traced back to answers from the outdated page. An oracle scores that retrieval as a hit. So I add a third number, the version share: for each case the oracle gets right, how many documents in the corpus answer the same question differently.
Why should a golden set be picked by hand and not sampled?+
Because the cases that do the most harm are rare. The course's first version drew 40 questions at random and none landed on a claim whose reserve had never been set, which is roughly one claim in seven in its corpus. The suite could not tell a system that says not yet established from one that invents a number. Two cases were then added on purpose, including a claim that does not exist, where the only right answer is a refusal.
Where does this go in a forward deployed engagement?+
In week one, before the first feature, and then in the handback packet. The packet I describe in The FDE Handback carries the customer's eval rows. The three shares belong on the same page, because they tell the product team whether the next fix is a retrieval change, a model change, or a content cleanup the customer has to own.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn