Your approval step is a formality by Thursday
AI prepares, a human reviews and approves.
That’s kind of the current, standard starting spot for AI automations in small businesses.
It feels really good too: you’re not putting anyone out of a job, AI isn’t “taking over,” and at the end of the day there’s a human to be held accountable for whatever action, in whatever particular business.
There is a slight problem with this though, and we want to dive into it today.
The finding is forty-three years old
In 1983, Lisanne Bainbridge published “Ironies of Automation” in Automatica (Vol. 19, No. 6, pp. 775-779). She was a psychologist at University College London, writing about power networks and chemical plants. Every page of it lands on your invoice queue.
Her argument: automate most of a task and the human you left in charge of the rest gets worse at it. Two mechanisms, both mechanical. Skills rot without practice: “physical skills deteriorate when they are not used,” so “a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one.” And attention has a shelf life. Citing Mackworth’s 1950 vigilance research, Bainbridge writes that it is “impossible for even a highly motivated human being to maintain effective visual attention towards a source of information on which very little happens, for more than about half an hour.”
Half an hour. Then the line every owner should read twice: “the operator will not monitor the automatics effectively if they have been operating acceptably for a long period.”
The system running well is the thing that breaks the person watching it.
Anthropic published the receipt this month
A couple weeks ago, TechCrunch reported that Anthropic is turning Claude Code’s auto mode on by default for Pro, Max and Team plans. The company’s own numbers made the case. In a study of 1,053 paid testers, automated screening caught 89% of harmful actions. Human review caught 13.6%. And, reported as a general statement about Claude Code users rather than as a result from those 1,053 testers, people approve 97% of permission prompts.
The screen caught roughly six and a half times what the humans caught, and the humans here were paid developers reviewing changes to their own code. The explanation TechCrunch offers, quoting Anthropic’s own line: “manual review can become habitual.”
That is Bainbridge, in production, with a sample size.
Picture a 40-person plumbing contractor
Hypothetical, somewhere outside Richmond. Eleven trucks, four people in dispatch, one office manager who owns accounts payable, and an owner who signs anything over $500. They wire up an AI that pulls supplier invoices out of email, matches each one to a purchase order and a job number, and stacks them in a queue. The owner sets the rule every owner reaches for: the software can prepare it, I approve it.
Monday he reads all thirty. He catches a duplicate freight charge and feels good about the rule.
Tuesday he reads most of them.
Wednesday he checks the vendor and the total.
Thursday it’s a keystroke. Two weeks in, he clears the queue from the truck at a red light, and the approval step has become a place where his name goes.
Nothing went wrong. That is the mechanism. Thirty correct invoices in a row are exactly the training required to stop reading the thirty-first.
The ten minutes of control you actually bought
The Census Bureau reported on 2026-08-11 that 55% of U.S. workers have used AI on the job for at least one of 11 tasks. Among those who used it in the last week, 25% saved less than an hour a week. That is the honest baseline for most AI at work right now: a person, a chat box, forty minutes.
Run the same arithmetic on the approval queue. Thirty invoices a week at 20 seconds each is ten minutes of human attention standing between the shop and every dollar leaving AP. Bainbridge’s half hour of vigilance, stretched across a week, spent on a screen where nothing ever happens. Operators buy that ten minutes and call it a control, and it is the pattern we keep watching.
Build the refusal
The design that survives contact with a real week has three properties.
- It refuses when it cannot know. Invoice matches the PO within tolerance, vendor on file, job number valid: post it, log it, no prompt. Freight line 40% above that vendor’s last three, no PO on record, or a total past the threshold: stop, and hold it.
- It names the reason it stopped. “Held: freight $412 against a $290 trailing average, vendor Southside Supply, job 4471.” An owner rules on that in eight seconds from a phone. “Approve this invoice?” teaches him nothing and earns a yes.
- It interrupts a human only where a human decides better. Matching and arithmetic go to the machine, which is better at both and never gets tired on a Thursday. Changed bank details on a known vendor go to a person every time, because that call is about trust and no rule holds it.
Count the interrupts, because the count is the whole design. A system that stops the owner twice a week gets read twice a week. A system that stops him thirty times gets a keystroke. Most shops set that number by accident, at “all of them,” and then wonder why the person in the loop stopped seeing anything.
The failure rhymes with the one in Bolt a motor to a line shaft and you’ve still got a line shaft: the AI is new, the approval ritual is inherited whole from the paper workflow, and nobody rebuilt the shape of the work underneath it.
By Thursday
An approval step that approves 97% of what it sees is a record of attendance. It documents that a human sat near the keyboard while the money moved, it satisfies whoever asks who signed off, and in week six it will pass the duplicate freight charge straight through.
Bainbridge could have told you that in 1983. She did.