Does keeping a human in the loop make an AI agent safe?

THE SHORT ANSWER

Not as a sentence in a safety case. Jakob Nielsen's 2026 review of vigilance research shows two failures. A person watching for rare events detects 10 to 15% fewer within 30 minutes. A person answering approval prompts habituates, so the request that deserved a no is approved with the rest; in a study by Ting Yan that he reports, users who wrote their own ask-me-first rules approved 67% of the resulting prompts and blocked less overreach than users who judged every action by hand. A human gate works when it is reserved for irreversible, external actions, shows the context for the decision, and is measured. Keep a gate ledger: one row per place a human is supposed to catch something, with asks per week, approval rate, median seconds to approve, and the date of the last no. Nielsen's thresholds for theater are an approval rate above 95% or a median decision under 2 seconds. For gates replaced by undo, record the undo window next to the gap between the user's sessions, because a window that closes while the user is away protects nothing.

The phrase promises a person who catches the bad action. The research Jakob Nielsen collects says that person stops catching things on a schedule: within half an hour when watching, and within a few dozen routine approvals when answering prompts. Neither failure is about effort, and telling people to pay attention has never fixed either.

What works is fewer, better gates and a count of what each one does. Reserve the interrupt for actions that can't be taken back and that leave the account. Give reversible actions an undo with a window the user will be present for. Then measure every gate, because the only evidence that a human is in the loop is a refusal.

The argument, the ledger columns, and where undo stops helping are in Kill 'Human in the Loop' as a Control. Keep a Gate Ledger. The design of the undo itself is in Undo Is a Design Primitive.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-09 · 1 min read