Call Join free
Tech Lab Miami
Tech Lab MiamiTECH LAB MIAMI
← Back to Automation & Operations

Where to Put the Human in an Automated Workflow

Discover effective strategies for integrating human oversight into automated workflows to enhance quality and decision-making. Learn about the importance o
Where to Put the Human in an Automated Workflow

Rite Aid had humans in the loop. It didn't help.

Between 2012 and 2020, a national pharmacy chain ran facial recognition across hundreds of stores to flag suspected shoplifters. A human always made the final call: an employee got the alert, walked the floor, and decided what to do. That is textbook human-in-the-loop design.

It still ended in a federal enforcement action. According to the FTC's December 2023 announcement, the system produced thousands of false-positive matches, and employees acting on those alerts followed customers, searched them, ordered them out, and called police. The FTC's complaint listed what was missing around the human step: no accuracy testing before deployment, no tracking of false-positive rates or of the actions taken on them, no adequate training that the system could be wrong, and — after the company adopted a tool that let staff report a "bad match" — no follow-through to ensure anyone actually used it. The settlement banned facial recognition for surveillance for five years.

That's the operational lesson worth paying for: a human in the loop is not a control. A human with context, authority, time, and a measured error rate is a control. Everything else is a signature on someone else's output.

Why "add a reviewer" quietly stops working

The failure mode has a name and a research base. In their review of complacency and automation bias in Human Factors, Parasuraman and Manzey found that operators stop scrutinizing automation that has been consistently right — and the effect is measurable. In one multitask experiment they analyze, participants caught roughly a third of automation failures when the system's reliability stayed constant, versus around 82% when reliability visibly varied (Parasuraman & Manzey, 2010). Complacency showed up in novices and experts alike, and got worse under competing task load.

Read that backwards and it stings: the better your automation gets, the weaker your human review becomes. A reviewer approving 400 correct outputs in a row is being trained, every day, to click approve. By the time the system produces the output that matters, the reviewer is a rubber stamp with a payroll number.

Regulators have started writing this into law. The EU AI Act names automation bias directly: Article 14 requires that high-risk systems be built so people assigned to oversight can understand the system's limits, "remain aware" of the tendency to over-rely on its output, and disregard, override, or reverse a decision (Regulation (EU) 2024/1689). Separately, GDPR Article 22 gives people subject to solely automated decisions with legal or similarly significant effects the right to obtain human intervention (Regulation (EU) 2016/679). If you sell into the EU, oversight isn't a design preference. If you don't, it's still the standard your enterprise buyers will use to score you.

Four places the human can sit — pick deliberately

"Where does the human go" has exactly four answers. Most teams default to the second one and never revisit it.

  1. In front (gate): nothing executes until a person approves. Highest safety, lowest throughput. Reserve it for actions you cannot walk back.
  2. Inside (escalation): the workflow runs itself and routes only defined exceptions to a person — low confidence, high value, unusual counterparty, missing data. This is where most operators should live.
  3. Behind (sampling and audit): everything ships, and a person reviews a sample plus every complaint, reversal, and refund. Cheap, and the only placement that keeps working at volume.
  4. Around it (design and rules): the human sets thresholds, writes the escalation rules, reads the error reports, and decides when to shut the thing off. This one is never optional and is almost never staffed.

The test that picks the placement for you

Skip the philosophy. Score each automated action on three axes:

  • Reversibility: can you undo it in an hour, or is it in a customer's inbox, a public feed, a bank ledger, or a legal record?
  • Blast radius: does an error hit one record or ten thousand? Batch jobs deserve gates that single actions don't.
  • Detectability: if this goes wrong, does something break loudly, or does it silently produce plausible garbage nobody catches for a quarter?

Irreversible plus wide plus quiet gets a gate. Reversible plus narrow plus loud gets post-hoc sampling. Everything in between gets escalation rules — and the rules, not a person's vigilance, are what carry the load.

Note the trap in the third axis. Generative outputs — drafted replies, summaries, extracted fields, generated copy — score badly on detectability precisely because they read fluently when they're wrong. Fluency is not accuracy, and a reviewer skimming a well-formatted paragraph is checking tone, not facts. If you're routing model output into a customer-facing surface, either narrow the task until errors are checkable in seconds, or gate it. Teams working through this trade-off on live systems is a good chunk of what we do in AI Consulting, and it's the same triage we apply when scoping business automation work.

Route a slice, not the flood

Look at how mature platforms ship this. Stripe's fraud tooling doesn't ask a human to look at every payment — its documentation frames reviews as a way to supplement automated systems with human expertise, using rules you write to place a targeted subset into a prioritized queue: elevated risk score, out-of-country, above an amount threshold, unusual email domain (Stripe Radar documentation).

Three things are worth stealing from that pattern:

  • The queue is a product surface, not a spreadsheet. The reviewer sees the risk signal and the surrounding context in one place.
  • The routing criteria are explicit and editable. You can tighten or loosen them as your error data comes in.
  • The reviewer has real verbs. Approve, refund, mark as fraudulent. Not "acknowledge."

If your review step is a Slack message saying "the automation ran, take a look," you have built a notification, not a control. Any decent workflow tool can do better; the constraint is design, not tooling. We break down these queue patterns in more depth in Learn.

Give the reviewer something to disagree with

A reviewer who can only approve or approve-later will approve. Build the disagreement path first:

  • A one-click reject that costs the reviewer nothing socially. If rejecting means a meeting, rejections go to zero.
  • A visible confidence or risk signal, so review effort scales with uncertainty instead of spreading evenly across everything.
  • A "this was wrong" report on outputs that already shipped, and — the part the FTC case turned on — active enforcement that people use it.
  • A documented owner. The NIST AI Risk Management Framework's Govern function is built around exactly this: defined roles, accountability structures, and oversight of human-AI configurations, with training tailored to the people in oversight roles (NIST AI RMF 1.0; see also the AI RMF Playbook). "Whoever is on shift" is not an owner.

Measure the human step or you don't have one

You already instrument your automation. Instrument the person too — not to police them, but because these numbers are the only evidence your control is alive:

  • Override rate. If it's near zero for weeks, either your routing is too loose or your reviewer has checked out. Both are findings.
  • Queue depth and dwell time. Volume is the enemy of attention. A queue that grows faster than it drains converts review into rubber-stamping on a predictable schedule.
  • Sampled defect rate on auto-approved output. The number that tells you whether the gate you removed should go back.
  • Downstream reversals: refunds, retractions, corrections, complaints. These are your ground truth, and they arrive late — which is why you track them separately from anything the reviewer self-reports.

Set a review cadence on these before launch, and write down the threshold that triggers a rollback. The decision to keep running a degraded system is much harder to make in the moment than in advance.

What to do this quarter

Take your highest-volume automated workflow and do four things. Map every action it takes that touches a customer, a payment, or a public surface. Score each one on reversibility, blast radius, and detectability. Move the human from wherever they happen to be sitting to wherever the score says they belong — usually out of the approval seat and into the exception queue plus a weekly sample. Then put a name, a metric, and a review date on it.

That's a week of work, and it's the difference between automation that compounds and automation that quietly accumulates liability. If you're building or replatforming the systems these workflows run on, the same questions belong in the architecture from day one — see Software. And if you'd rather pressure-test your design against other operators before you ship it, that's what the community is for.

FAQs

What is human-in-the-loop automation?

It's any workflow where a person can inspect, approve, override, or reverse what an automated system does. The label is less useful than the specifics: which decisions reach the person, what they can see when they get there, and what authority they hold. A workflow where a human clicks approve on everything and can see nothing is human-in-the-loop in name only.

Doesn't adding human review just slow everything down?

Blanket review does. Targeted review usually doesn't, because you're routing a defined slice — low confidence, high value, unusual pattern — rather than the whole stream. The real cost isn't latency, it's attention: every item you route consumes reviewer focus that the genuinely risky items need.

How do I know if my human review step has gone stale?

Check the override rate and the queue dwell time. A reviewer approving nearly everything, or a queue draining slower than it fills, are the two visible symptoms of automation complacency — a pattern the human factors research associates with highly reliable systems and heavy task load, not with careless people.

Am I legally required to keep a human in the loop?

It depends on jurisdiction and use case, and this isn't legal advice. Two anchors are worth reading directly: the EU AI Act's Article 14 sets human oversight requirements for high-risk systems, and GDPR Article 22 gives people a right to human intervention for solely automated decisions with legal or similarly significant effects. In the US, the FTC has brought enforcement under existing consumer protection authority against automated systems deployed without reasonable safeguards.

Where should a small team start?

With the irreversible actions. Anything that sends money, publishes publicly, deletes data, or contacts a customer under your name gets a gate until you have error data proving it doesn't need one. Everything else can start with sampling. You can find upcoming sessions on this and related operational topics under Events.

Visuals

Editorial photograph illustrating human review inside an automated workflow: colleagues examining a decision together in a working environment.
Human review only functions as a control when the reviewer has the context to judge, the time to look, and the authority to say no.
← Back to Automation & Operations
Chat with us