ExecutionNew·Falk Gottlob··6 min read

The Menu Is Not the Error Analysis

Hamel Husain tested Claude's auto eval tool and said hold off. Use it to find failures, then read fifty outputs and rank before any grader is built.

AI evalserror analysisauto evalClaude Codebuild_evalHamel HusainIsaac FlathShreya ShankarTeresa TorresLLM as a judgebottom-up evalsthe rewrite
Helpful?

Execution Falkster cover on teal: a closed cream menu standing beside a tall stack of paper sheets, with a magnifying glass showing an orange square resting on top of the stack.

Anthropic shipped an auto eval tool inside Claude Code, and the first serious review of it landed this week. Hamel Husain tested the new commands on camera with Isaac Flath and concluded he would hold off for now. I would not hold off. I would use it this week for the half that works, and I would close the menu it opens with.

The short version

Claude's auto eval tool, the build_eval and hill-climb commands in Anthropic's claude-api plugin, is good at finding candidate failures and too quick to turn one into an eval. Hamel Husain's review says it asked him to pick an eval from a recommended menu before he had read the data, asked for label sign-off without context, and bundled four failures into one evaluator. He also calls its discovery the strongest one-shot issue finding he has seen. That corrects what I wrote in August about models and bottom-up evals. Two things did not change: ranking and slicing. So "build the eval" becomes four steps. Take the tool's list. Close the menu and read fifty outputs. Rank by count times customer cost. Then one failure per grader, with the judge prompt read by a person.

What Hamel found

His post is short and specific, and you should read it. The setup was conversation traces from an apartment leasing assistant.

Claude began by suggesting several potential failures and asked them to pick one to turn into an eval straight away. It recommended call-transfer rules. They had not reviewed the conversations yet, so they had no way to know whether that was a real failure or worth going after first. They took the recommendation anyway, on the reasoning that a typical user would.

Next it wrote markdown files and asked them to skim the inputs and say which labels were wrong, in chat. They ended up asking Claude to build a small web app for annotation and used that. Later it showed aggregate label counts and asked whether they would have scored any case differently, without showing enough to answer.

Then the evaluator. It checked four things at once. Whether the assistant asked for confirmation the right number of times, which needs an LLM judge. And three things code can check: talking between consent and transfer, talking during or after the transfer, and saying tool mechanics out loud. A case passed only if all four passed.

And the good part. Hamel says the tool's ability to discover issues out of the box beat other auto-eval approaches he has tried, across human handoff, formatting, and voice. He has since talked to the plugin's author, who took the feedback and plans changes.

The part I had wrong

In August I wrote The Wince Is the Spec: Bottom-Up Evals a Model Can't Write. The argument, built on Shreya Shankar's distinction, was that a model drafts top-down evals well, because they are the brief in a new shape, and is bad at the bottom-up ones that come from reading outputs.

Hamel's report says the reading has improved. A tool that looks at traces and comes back with handoff problems and voice problems is doing some of the bottom-up work. I would write that August post differently today. The model can now have a decent first look.

What I do not know is whether it would have found my failure. On Heidi, the action that bothered me was permitted, reversible, logged, and in budget. It was wrong because the evidence behind it was thin. No rule named that. Maybe a discovery pass lists it and maybe not. I have not run the plugin on those traces, so I will not claim either.

Two things that did not move

Ranking. A menu with a recommended item is the tool's guess at what matters. It has no idea what a failure costs your customer. The error analysis log in The AI Eval Starter Kit: Six Templates, Error Analysis to Gate is fifty rows, one note each, then a tally ranked by count times customer cost. The top three errors get an eval this cycle and nothing else does. The one discipline I ask for is not opening the judge template until that log is full. The menu skips the log.

Slicing. In The Eval That Caught What the Demo Did Not I ran a feature against 200 real cases. The average was fine. The long-input slice, about a fifth of our real traffic, failed badly, with clean and confident wrong output. The average hid it.

A flag that passes only when four checks pass is the same blindfold from the other side. When it goes red you cannot tell which of the four fired. When it stays green on a sample you cannot tell which of the four was never exercised. It also puts three cheap code checks behind one expensive judge call. Hamel says the same thing more politely: one error at a time, or at least split code from judge.

The rewrite

"Build the eval" becomes four steps, in this order.

  1. Run the discovery pass and take the list. Treat it as candidates. This is the part the tool is good at.
  2. Close the menu. Read fifty real outputs yourself, one note each. If reading in a text file is painful, have the agent build the annotation page first, as Hamel and Isaac did.
  3. Rank by count times customer cost. Top three only.
  4. One failure per grader. Code where code can decide it, a judge only where it cannot, and a person reads the judge prompt before anyone trusts the score.

Then let hill-climb loose on graders you understand.

What to do this week

If someone on your team has installed the plugin, send them the four steps before they accept the first recommendation. If nobody has, fifty rows in a spreadsheet still works and costs an afternoon.

The plugin will change. Its author said so. The order should not, because it is the part of the job that did not get cheap. That is the argument on AI Product Management: AI collapsed the cost of the parts PMs were trained to be good at and left the parts nobody trained for. Deciding which failure matters is one of those.

Related answer: Should an auto eval tool pick your first eval?

Sources: Claude's new auto eval tool, Hamel Husain, September 30, 2026. AI Evals: A Hands-On Guide for Product Teams, Teresa Torres, Product Talk, September 2, 2026. The Wince Is the Spec: Bottom-Up Evals a Model Can't Write, falkster.com, August 24, 2026. The Eval That Caught What the Demo Did Not, falkster.com, August 12, 2026.

Share this post

Frequently asked

What is Claude's auto eval tool?+

Per Hamel Husain's review of September 30, 2026, Anthropic's claude-api plugin for Claude Code now includes a build_eval command and a hill-climb command that help you build evals, check the graders, and improve an application against them. He and Isaac Flath tested it on a livestream using conversation traces from an apartment leasing assistant.

What did Hamel Husain find wrong with it?+

Three things. It pushed them to pick an eval from a menu of suggested failures before they had looked at the data. It asked them to validate labels without enough context, first in a markdown file and later from aggregate label counts. And the evaluator it built bundled four different failures, one needing an LLM judge and three checkable in code, behind a single flag that passes only if all four pass. His verdict was to hold off for now, and he notes the plugin's author plans to update it.

What did it do well?+

Discovery. Hamel writes that it was the strongest performance he has seen from a one-shot issue discovery approach, finding problems with human handoff, formatting, voice agents, and more. That is better than I gave models credit for in August, when I wrote that the bottom-up half of evals is the half a model cannot do.

Should a team use the tool now?+

I would, for the discovery pass and for the checks that can be written from the task description alone. I would not pick from its menu. Take the list of candidate failures, read fifty real outputs by hand, rank what you find by count times customer cost, and build an eval only for the top three. Then one failure per grader, and read the judge prompt before trusting it.

Why not bundle several failures into one evaluator?+

Because a single pass flag hides which failure fired. I had a feature score well on average across 200 real cases while one slice, about a fifth of our traffic, failed badly. The average hid the slice. An all-must-pass flag does the same thing from the other direction, and it also mixes cheap code checks with an expensive judge call.

What is error analysis in evals?+

Reading real outputs by hand and writing down what is wrong with each one, before building any automated check. In my starter kit it is fifty rows, one note each, then a tally ranked by how often an error occurs times what it costs the customer. The method is Teresa Torres's first step, and Hamel's review ends on the same point: a tool that does not put looking at data at the center is not worth using.

THE SHORT ANSWER

PART OF

Enterprise AI Agents

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.