Should an auto eval tool pick your first eval?

THE SHORT ANSWER

No. Use the tool to find candidate failures and do the picking yourself. Hamel Husain's review of Claude's auto eval tool found that it recommended an eval from a menu before he had read any data, and also that its one-shot issue discovery was the strongest he has seen. So take the list and close the menu. Read fifty real outputs by hand, one note each, rank the errors by count times customer cost, and build an eval for the top three only. Then write one grader per failure. I had a feature score well on average across 200 real cases while a slice worth about a fifth of our traffic failed badly, and a single flag that bundles four checks hides a failure the same way an average does.

Anthropic's claude-api plugin can now propose and build evals inside Claude Code, and The Menu Is Not the Error Analysis is my read of Hamel Husain's review of it: keep the discovery, skip the menu. The fifty-row log I use for the ranking step is in The AI Eval Starter Kit: Six Templates, Error Analysis to Gate, and the case for breaking results into slices is The Eval That Caught What the Demo Did Not.

It belongs to the AI Product Management argument: deciding what correct means, and which failure matters, is the part of the job that did not get cheap.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-01 · 1 min read