Anthropic's claude-api plugin can now propose and build evals inside Claude Code, and The Menu Is Not the Error Analysis is my read of Hamel Husain's review of it: keep the discovery, skip the menu. The fifty-row log I use for the ranking step is in The AI Eval Starter Kit: Six Templates, Error Analysis to Gate, and the case for breaking results into slices is The Eval That Caught What the Demo Did Not.
It belongs to the AI Product Management argument: deciding what correct means, and which failure matters, is the part of the job that did not get cheap.