Logo home 9

Over 10 years we helping companies reach their financial and branding goals. Onum is a values-driven SEO agency dedicated.

CONTACTS
AI Automation AI / ML Development

Human-in-the-Loop AI Automation: Where to Put the Person

Category
AI Automation
Read Time
9 min read
Published
October 7, 2026
Status
Published

Every AI automation has a person in it somewhere. The design question is where. The four oversight patterns, a worked threshold calculation for 10,000 invoices a month, and how to stop reviewers rubber-stamping.

Almost every AI automation project ends up with a person somewhere in the loop. The question that decides whether the project works is not whether to include one, but where. Put the person at every step and you have built an expensive way of doing the work twice. Put them nowhere and the first confident mistake reaches a customer, a supplier or the general ledger before anyone notices.

This guide to human-in-the-loop AI automation covers the four places a person can sit, how to choose a confidence threshold with arithmetic rather than instinct, and the problem most designs ignore: reviewers who stop reading because the machine is usually right.

Patterns

Four Places to Put the Person

Human-in-the-loop is often described as a single design. In practice there are four distinct patterns, and most production systems combine three of them.

PatternHow it worksUse it forCost
Approve before actingThe system prepares the action; a person approves every oneIrreversible or high-value actions: payments, contract changes, deletionsSlow, and reviewers tire quickly
Exception-only reviewCases below a confidence threshold go to a person; the rest proceedHigh-volume work with a clear right answerDepends entirely on where the threshold sits
Act, then sampleEverything proceeds; a person audits a random sample afterwardsLow-impact, easily reversed actions such as tagging and routingCheap, but errors are found after the fact
Escalate by ruleFixed triggers send a case to a person regardless of confidenceKnown danger: new suppliers, changed bank details, values above a limitLow, and the most underused of the four

The usual production design is exception-only review for the bulk of the volume, escalation rules for the cases where being wrong is expensive, and a small audit sample of everything that went through automatically. The audit sample is what tells you whether the threshold is still right six months later.

Confidence tells you how sure the model is. It does not tell you how much a mistake costs. Escalation rules exist for the second question.

Thresholds

Choosing a Threshold With Numbers

Most teams pick a confidence threshold of 0.9 or 0.95 because it sounds safe. The right threshold is the one where the cost of reviewing cases and the cost of errors that slip through add up to the smallest total. You can find it from a labelled sample of past work.

Here is a worked example for a finance team processing 10,000 supplier invoices a month. The team ran the model over 2,000 historical invoices whose correct values were already known, then measured what would have happened at four thresholds. A review takes two minutes at a loaded cost of $30 an hour, and an error that reaches the ledger costs about $40 to find and correct.

ThresholdProcessed automaticallyError rate in thoseReview costError costMonthly total
0.9852%0.2%$4,800$420$5,220
0.9571%0.5%$2,900$1,420$4,320
0.9083%1.2%$1,700$3,980$5,680
0.8091%3.0%$900$10,920$11,820

At 0.95, 2,900 invoices go to a person (97 hours, $2,900) and about 36 errors slip through ($1,420). That is the cheapest row. The highest automation rate, 91 percent, is the most expensive option by a wide margin, because its errors cost more than all the review time it saves.

Now change one number

Suppose a single error can send a payment to the wrong bank account, and each one costs $400 rather than $40. At 0.95 the errors now cost $14,200 and the total jumps to $17,100. At 0.98 the total is $8,960. The cheapest threshold moves, because the threshold depends on the cost of being wrong, not on the model.

This is why one threshold rarely fits a whole workflow. A sensible design runs routine invoices at 0.95 and sends anything touching bank details to a person every time, whatever the model’s confidence. That second part is an escalation rule, and it is cheap because those cases are rare.

Rubber-Stamping

When Reviewers Stop Reviewing

A reviewer who approves 98 cases in a row soon learns to approve the 99th without reading it. This is automation bias, and it quietly turns a human check into a human signature. The review queue still exists and the audit log still shows an approver, but nobody is actually looking.

Show the evidence, not just the answer
Put the source document beside the extracted values, with the relevant line highlighted. A reviewer who has to hunt for the evidence will eventually stop hunting.
Say why it was flagged
“Low confidence on the total” or “new supplier” tells the reviewer where to look. A case that arrives with no reason gets a glance at everything and a proper check of nothing.
Ask for the decision, not a click
For high-stakes fields, have the reviewer confirm the value itself, for example by re-entering the last four digits of an account number. Confirming is work; clicking Approve is not.
Seed known errors
Insert a small number of cases with a deliberate, known mistake. The share that reviewers catch is the best available measure of whether the review is real.
Watch time per review
If the median review takes three seconds and the override rate is near zero, the queue is being cleared, not reviewed. Both numbers are easy to log and rarely are.
Feed corrections back
Every override is a labelled example. Collect them, add them to the evaluation set, and the next threshold decision is based on recent work rather than last year’s sample.
Regulation

Where Oversight Is Required, Not Optional

For most back-office automation, human review is a design choice made on cost and risk. In some areas the law makes it for you.

Under GDPR Article 22, people have the right not to be subject to a decision based solely on automated processing when it has legal or similarly significant effects on them, such as a refused credit application. If a workflow makes decisions about individuals, a meaningful human step is often the simplest way to stay on the right side of that rule.

The EU AI Act goes further for systems it classes as high-risk, which include AI used in hiring, credit scoring and access to essential services. Article 14 requires those systems to be designed so a person can understand the output, intervene, override it and stop the system. The 2026 Digital Omnibus pushed the main high-risk deadline back to December 2027, but it did not change what Article 14 requires. Most invoice, ticket and document workflows fall outside the high-risk categories, but anything that ranks applicants or decides eligibility should be checked early rather than retrofitted.

Design

Building the Review Step Properly

1Decide the cost of an error first
Write down what a mistake costs for each kind of case before choosing any threshold. Without that number, every threshold is a guess dressed up as caution.
2Measure on your own labelled data
A few hundred to a few thousand past cases with known answers is enough to draw the table above. Vendor accuracy figures describe someone else’s documents.
3Write escalation rules for known dangers
Bank detail changes, first-time suppliers, values above a limit, legal wording. Rules are cheap, predictable and easy to explain to an auditor.
4Give the queue an owner and a deadline
Exceptions that age unseen are worse than no automation, because everyone assumes they were handled. Track queue age and alert on it.
5Audit a sample of the automatic path
Two to five percent of automatically processed cases, checked every month, shows whether the error rate is drifting. It is the only window onto the cases nobody looked at.
6Revisit the threshold on a schedule
Suppliers, formats and models change. Re-run the threshold table quarterly with the corrections collected since the last run.

Getting this layer right is also what makes the business case hold. The exception path is where most of the remaining cost lives, which is why it features so heavily in AI automation ROI and in the economics of AI document processing automation. For agents that can take actions rather than just suggest them, the approval step is also a security control, covered in AI agent security.

FAQ

Frequently Asked Questions

What is human-in-the-loop AI automation?
It is automation designed so that a person reviews, approves or corrects some of the system’s work before or after it takes effect. The person might approve every action, review only low-confidence cases, handle cases that trigger fixed rules, or audit a sample afterwards.
What confidence threshold should I use?
The one where review cost plus error cost is lowest for your process. Measure it on a labelled sample of past cases. In the worked example above that was 0.95, but raising the cost of a single error from $40 to $400 moved the best choice to 0.98.
How do you stop reviewers rubber-stamping?
Show the source evidence beside the answer, state why each case was flagged, ask for confirmation of key values rather than a single click, seed occasional known errors to measure attention, and track time per review and override rate.
Does human review remove the benefit of automation?
Not if it is placed well. In a typical high-volume workflow, most cases proceed automatically and a person sees only the uncertain or risky minority. The saving comes from not handling routine cases, not from eliminating the person.
Is human oversight a legal requirement?
Sometimes. GDPR Article 22 restricts solely automated decisions with legal or similarly significant effects on individuals, and the EU AI Act requires human oversight for high-risk systems such as those used in hiring or credit scoring. Most back-office workflows are not in those categories.
How much of the automatic output should be audited?
A random two to five percent each month is a practical starting point. It is enough to detect a drifting error rate without recreating the manual workload, and it should rise temporarily after any model or process change.
Wrapping Up

Conclusion

The person in the loop is not a safety blanket thrown over the automation at the end. It is a component with a cost, a failure mode and a position that should be chosen deliberately. Decide what an error costs, measure where the threshold should sit, write rules for the dangers you already know about, and check that the reviewers are still reading.

Most failed automation projects did not lack a human step; they had one in the wrong place or one that had quietly stopped working, as covered in why AI automation projects fail. We design AI automation with the review step treated as part of the system, measured like the rest of it.

About MetaDesk Global

Engineering the Next Generation of Connected Products

MetaDesk Global helps startups and enterprises develop intelligent connected products that combine embedded systems, Industrial IoT, Edge AI, and cloud technologies. Our expertise includes:

Industrial IoT (IIoT) Solutions Embedded Firmware Development Edge AI Development Predictive Maintenance Systems PCB Design IoT Gateway Development Cloud Integration OTA Firmware Updates AIoT Product Development End-to-End Product Engineering

From hardware design to AI-powered industrial platforms, we build scalable solutions for the next generation of connected products.

Start Your Project

Building a Connected Product?

We design IIoT sensor networks, Edge AI pipelines, and secure cloud platforms — from prototype to production.

Request a Free Quote →