How much of an AI's reasoning should you show the user?

THE SHORT ANSWER

As much as changes what a reasonable reader would do next, and no more. Use a five-level ladder, from the answer alone up to the raw trace, and decide which rung is the default for which reader in which moment and what single action moves them up. Showing everything is not transparency, it is a way of making the user responsible for checking work they never asked to supervise. The highest-value part is the escalation trigger list, the conditions under which the product raises the level on its own.

Most AI products are sitting on one of two failures.

A confident paragraph with nothing behind it, delivered in the same voice whether the system checked forty documents or guessed. Or the overcorrection: forty visible reasoning steps, a retrieval log, and an animation of the model thinking, all stacked above the answer the person came for.

Same mistake. Nobody decided what the reader was supposed to do with the information.

Five rungs on the ladder. Level 0 is the answer alone. Level 1 adds provenance, what it drew from, one line and one action to open. Level 2 shows the shape of the reasoning, two or three steps in the user's language rather than a transcript. Level 3 is the full chain on request, available in one action and never a default. Level 4 is the raw trace: prompts, tool calls, tokens, timings, model versions.

Drawing the ladder is the easy part. Design work lives in two columns: where each reader starts, and what sits one action away.

A first-time user starts at 1 with 2 one action away. A routine user on a low-stakes task sits at 0 or 1. A routine user on a high-stakes task starts at 1 with the full chain reachable. A reviewer or approver starts at 2. An admin or auditor starts at 2 with the raw trace available. Support and engineering start at 3.

Level 4 is never a default for anyone. It is a debugging tool that occasionally gets mistaken for an interface.

Now the test, short enough to use in a review. Does this disclosure change what a reasonable reader would do next? If not, it is decoration, and decoration in this position is not free. It spends the attention that the disclosures which do matter will need.

Four cases where more costs you. All four are specific enough to check against your own product this afternoon.

The trace reveals the method is thinner than the tone implied. A three-sentence answer delivered with authority whose trace shows one document read worse than the answer alone. Fix the tone, not the trace.

The volume implies rigor that is not there. Forty steps feels thorough and is often verbose, and users read length as effort, which is the wrong signal from a system where length is free.

It exposes the user's own data back to them unexpectedly. "I read your last two hundred messages to answer this" can be true, accurate, and a serious mistake to say in that moment.

The reader cannot evaluate it. A chain of reasoning shown to somebody with no way to judge it creates anxiety with extra steps.

The most valuable page in the template is not the ladder, it is the escalation list: the conditions under which the product raises its own default without being asked. Confidence below your stated floor. An action in an irreversible class. Output that contradicts something the user or the system said earlier. A user who reversed a similar output recently. Money, permissions, external communication, or deletion involved.

Write those triggers down and put them in the eval set. A ladder that only moves when the user climbs it is a ladder nobody climbs on the day it matters, because the day it matters is the day they are moving fast and trusting the product.

Escalation is also what makes a low default defensible. You can argue for Level 0 on routine output precisely because the system raises itself when the situation changes. Without triggers, the argument collapses into a fight between the person who wants a clean screen and the person with the scariest hypothetical, and the hypothetical wins every time.

This week: find the surface where your product shows the most reasoning and ask three users what they do with it. If the answer is that they skip it, drop that surface a rung.

The five levels, the defaults table, the four failure cases, and a blank trigger list are in The Disclosure Ladder. The full argument is in Disclosure Without Overwhelm.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-07 · 4 min read