Human-in-the-loop (HITL) is a design pattern where a system cannot proceed past a certain point without explicit human input. The person reviews, approves, corrects or rejects, and the system continues based on that response.
Where the term comes from
The phrase originates in control systems and machine learning, where a human was needed to label data or validate model outputs. In modern software engineering, it has been adopted more broadly to describe any workflow where automation is deliberately interrupted for human review.
Why systems are built this way
Full automation is not always appropriate. The reasons for inserting a human step typically fall into a few categories:
- Risk: the consequence of a wrong action is high enough that the cost of a pause is justified. A payment above a threshold, a deployment to production, a deletion of records.
- Uncertainty: the system has low confidence in its output and flags this rather than guessing.
- Regulation: some domains require a human decision by law or policy. Credit decisions, medical recommendations, certain data processing operations.
- Edge cases: the input falls outside the range the system was designed to handle reliably.
The pattern accepts lower throughput in exchange for lower error rates on consequential decisions.
HITL in AI-assisted systems
As more production software incorporates model-generated outputs, whether suggestions, summaries, classifications or actions, HITL has become a standard architectural concern rather than an edge case.
An AI coding agent can propose a pull request, but a human reviews and merges it. A RAG pipeline can return an answer, but a support agent reads it before sending. An agentic workflow can execute a sequence of steps, but requires confirmation before side effects like sending emails or modifying a database.
The design question is where in the flow to place humans, and what information to give them when they arrive at that point, rather than whether to include them at all.
What makes a HITL step useful or useless
A human checkpoint only works if the person reviewing it can actually evaluate what is in front of them. This sounds obvious, but it fails in practice when:
- The output is too long or complex to read in the time available.
- The reviewer lacks the domain knowledge to spot errors.
- Approval has become a formality and people click through without reading.
- The system provides no explanation of how it reached its output, so the reviewer has nothing to work with.
When HITL becomes a rubber stamp, it provides the appearance of oversight without the substance. This is a meaningful engineering and organisational problem, not just a UX one.
Implementation patterns
Common approaches include:
- Approval queues: the system writes a proposed action to a queue; a human processes items in the queue before they execute.
- Confidence thresholds: the system routes high-confidence outputs automatically and flags low-confidence ones for review.
- Escalation paths: the system attempts resolution autonomously and escalates to a human only when it cannot proceed.
- Audit trails with rollback: the system acts immediately but logs everything and allows reversal within a window. This is HITL after the fact rather than before it.
The right pattern depends on latency tolerance, risk level and the volume of decisions the system handles.
The relationship to code review
Code review is itself a HITL mechanism applied to software changes. As AI-generated code becomes a larger share of what gets merged, the review step becomes more important, not less. The reviewer is the human in the loop for the coding agent's output. What this requires of engineers has changed somewhat; see what changed about code review for detail.
What we test for
HITL design reflects an engineer's judgement about where automation should stop and what a reviewer actually needs to see. Our vetting looks directly at this. In Session 1, prioritisation and technical understanding reveal whether someone can structure a problem before touching it. In Session 2, AI output verification and judgement by risk show whether they slow down appropriately when an agent produces something consequential, rather than approving it because it runs. See how we vet.
Short answers
Is human-in-the-loop the same as supervised learning?
Not exactly. Supervised learning uses human-labelled data to train models. HITL in engineering refers to runtime workflows where a person must act before a system proceeds. The concepts overlap in active learning pipelines but are otherwise distinct.
Does adding a human step make a system safer?
Only if the reviewer can meaningfully evaluate the decision. A checkpoint where people routinely approve without reading provides little protection. The quality of the human step matters as much as its presence.
When should a system remove a HITL step it previously had?
When there is evidence, usually from audit logs or error rates, that the human step is adding latency without catching errors. That evidence should come from measurement, not assumption.