
"Human in the loop" is usually one line in a safety case, a PRD, or a vendor's answer to a security questionnaire. There's no number beside it. Nobody can say how often the human said no.
Jakob Nielsen published the long version of why that line fails, and Kathy Baxter at Salesforce conceded the same week that the industry has no fix for it. I'd retire the line and replace it with something you can count.
The short version
Nielsen's article on vigilance splits human oversight into watch duty and verdict duty and shows both decay: detection of rare events falls 10 to 15% within 30 minutes of watching, and approval prompts turn into reflexes with each routine yes. His remedy for agents is to tier by consequence, prefer undo to ask, and instrument the rubber stamp. I agree with the direction and think undo gets too much credit, because an undo is a notification and a window, and both depend on the same attention that failed. The entry for the Kill List is the unmeasured sentence "a human reviews it." The replacement is a gate ledger: one row per human gate, with the action class, asks per week, approval rate, median seconds to approve, the last no, and for undo rows the window against the gap between sessions.
What Nielsen and Baxter said
Nielsen's piece is "User Vigilance Fails Twice in the AI Age," summarized in his October 2 roundup under the heading that vigilance fatigue dooms human-in-the-loop. I read the full article.
Watch duty first. Norman Mackworth's 1948 clock experiment had people watch a pointer for two hours and report rare double jumps. Detection fell 10 to 15% in the first 30 minutes and kept sliding. Urging people to try harder did nothing. Nielsen's point about AI review is uncomfortable: the better the system gets, the rarer the error, and the rarer the error, the worse a person is at catching it.
Verdict duty is the newer one. The agent works, and you get interrupted for a yes or no. He reports a 2026 study by Ting Yan in which 113 adults supervised an assistant through a simulated day. One group wrote standing permission rules in advance. They picked "ask me first" for 81% of their rules, then approved 67% of the prompts those rules produced. People judging each action fresh approved 40%. The rule writers blocked 40% of the overreach and the by-hand group blocked 60%. I haven't read the study, only his account of it. His line on it: "Feeling in control is not a safety metric."
His guidelines for agents: budget interruptions like money, tier by consequence, make every approval carry its context, and "prefer undo to ask." And a pair of thresholds I'm going to borrow. A median decision under 2 seconds, or an approval rate above 95%, means the gate is decoration.
Baxter, Salesforce's Principal Architect of Ethical AI Practice, published a summary of a white paper built from workshops with 21 organizations. It lists when an agent should escalate to a person (significant financial transactions, low-confidence decisions, detected prompt injection, repeated failure loops) and then says the quiet part: the value of agents is removing the human bottleneck, a platform that needs constant rescue isn't autonomous, and "the industry currently lacks a definitive solution to this tension."
Two posts, five days apart. One says the loop doesn't work the way the paperwork claims. The other says we don't know what replaces it.
Where undo runs out
"Prefer undo to ask" is the best single sentence in Nielsen's list, and I've argued a version of it in Undo Is a Design Primitive. It still carries the old problem inside it.
His wording for reversible actions is to act with a visible countdown and a one-click cancel. That's built for a person who is looking at the screen. Agents do a lot of their work when nobody is.
Undo for an action taken while the user slept is three separate things: the reversal, the notification, and the window. The notification is another interrupt. Send forty a day and people stop opening them, which is the rubber stamp running in the other direction. Call it the rubber ignore. And the window is where it quietly fails. If people open the product on weekday mornings and the undo window is four hours, the reversal exists in the code and not in anyone's life.
There's a class where undo is a comfortable lie. Reversible but witnessed: the database restores cleanly, and somebody already read the original message. State went back. The person who read it didn't.
So undo doesn't remove vigilance from the design. It moves it from the moment before the action to the hours after, and it needs measuring just as much.
The ledger
One row for every place a human is supposed to catch something. Approval prompts, review queues, exception readers, undo windows. If your architecture diagram has a little person icon on it, that's a row.
| Column | What goes in it |
|---|---|
| Gate | The action, in plain words |
| Class | Reversible private, reversible witnessed, irreversible contained, or irreversible external |
| Asks per week | Per person, not per system |
| Approval rate | Share answered yes |
| Median seconds to approve | From prompt shown to click |
| Last no | The date somebody refused |
| Window vs. gap | Undo rows only: the window, next to the typical time between that user's sessions |
Then read it for two failures.
A row over 95% approval or under 2 seconds is theater, per Nielsen. Remove it, re-tier it, or change what the prompt shows. I'd add the last-no column because the rate hides it: a gate nobody has refused in a quarter is guarding nothing, or it's being waved through, and you want to know which.
An undo row whose window is shorter than the gap is a different kind of theater. Lengthen it, or move the action up a class and ask.
The irreversible external rows are the ones that keep a person, in session, every time. I learned the money version the expensive way. My expense agent auto-approved anything under $500 with a matching receipt, a vague "dinner with customer" kept passing, and about $2,400 went out before anyone looked. That story is in Nothing Refunds a Sentence. So the ledger isn't a case for fewer humans. On those rows you want real refusals, and if the last no is blank, worry.
I should put one of my own recommendations through this. Three days before this post, in Fourteen Rules, and Who Blesses the Exception, I wrote that a person should read every exception a rule-watching agent blesses, counted by rule. I still think so. That reader is a gate, and I didn't give them a row. If they agree with 99% of blessings in under two seconds each, the reading has stopped. Count by rule, cap the volume one person sees, and plant a bad blessing now and then to check the reader is still reading. Nielsen's seeded-error advice is for exactly this.
What this changes in the seat
If you own an agent product, search your PRD and your security answers for "human in the loop," "human review," and "user approves." Each hit is a row you haven't filled in.
If you buy one, ask the vendor for the approval rate and median decision time on their riskiest gate, from a real customer. Most won't have it. That's the answer.
If you design the surface, the approval prompt and the undo notice are your screens, and they're graded by the two numbers above.
It sits under the argument on Enterprise AI Agents: what counts is agents still doing production work at ninety days and what each good outcome costs. A gate that approves everything is part of that cost. It just hasn't been billed yet.
This week: pick the one gate you'd least like to fail. Pull its last 100 decisions. Count the nos.
Related answer: Does keeping a human in the loop make an AI agent safe?
Sources: Jakob Nielsen, "User Vigilance Fails Twice in the AI Age: Watch Duty & Verdict Duty," UX Tigers, September 29, 2026, and "UX Roundup: Vigilance Fatigue," October 2, 2026; the Mackworth and Ting Yan findings are as he reports them. Kathy Baxter, "From Autonomy to Accountability: How to Think About Trust in the Multi-Agent Future," Salesforce News, October 7, 2026. Undo Is a Design Primitive, falkster.com, September 30, 2026.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
Does keeping a human in the loop make an AI agent safe?+
Not as a sentence in a safety case. Jakob Nielsen's review of vigilance research gives two reasons. A person watching a mostly empty stream detects 10 to 15% fewer rare signals within 30 minutes. A person answering approval prompts habituates, so the one request that deserved a no gets approved with the rest. A human gate is a control only when it is rare, carries its context, and is measured.
What is a gate ledger?+
One row for every place in an agent system where a human is supposed to catch something. Each row records the action class, how many asks reach the person per week, the approval rate, the median time to approve, and the date of the last refusal. Rows that use undo in place of approval record the undo window next to the typical gap between the user's sessions. It replaces the line human in the loop with numbers someone can audit.
What numbers show that an approval gate has become a rubber stamp?+
Nielsen's rules of thumb are a median decision time under 2 seconds or an approval rate above 95%. Either means the gate is approval theater and should be removed, re-tiered, or redesigned. I add one more signal: the date of the last no. A gate nobody has refused in a quarter is either guarding nothing or being waved through.
Is undo better than asking for approval?+
For reversible actions, yes, and Nielsen says so directly: prefer undo to ask. But undo relocates the attention problem. It has three parts, the reversal, the notification, and the window, and an agent that acts overnight delivers the notification to someone who is not there. A window shorter than the gap between a user's sessions protects nothing, and a flood of undo notices is ignored the way a flood of prompts is approved.
Which agent actions should always require a person?+
The irreversible external class from Undo Is a Design Primitive: money leaving, a message delivered outside the account, a permanent deletion, a permission granted, a write to a third-party system. I learned the money part from an expense agent that auto-approved anything under $500 with a matching receipt and let about $2,400 through. Those actions get a person in session every time, and that row of the ledger should show real refusals.
What did the Ting Yan study find about user-written permission rules?+
As Nielsen reports it, 113 US adults supervised an AI assistant through a simulated day of 18 actions, 7 of them unrequested. Participants who wrote standing rules chose ask me first for 81% of them and then approved 67% of the resulting prompts, against 40% for people judging every action fresh. The rule writers blocked 40% of overreach and the by-hand group 60%. I have not read the study itself, only Nielsen's account.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn