AI agents and business automation
Human-in-the-loop AI agent design that works in practice
Define where judgement, approval and escalation sit before an agent reaches live work.
Match oversight to consequence
Use lighter sampling for reversible low-risk tasks and explicit approval for actions affecting money, rights, customers or sensitive data. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. The review model should respond to harm and reversibility, not apply one blanket rule. Record the choice, the evidence still required and the person who will return with it. Keep the mechanism proportionate: the purpose is better judgement and follow-through, not additional ceremony.
Show the reviewer what matters
Present source material, confidence signals, changes made, policy checks and the reason an item was escalated. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. Review becomes theatre when a person sees only a polished output and no evidence. Record the choice, the evidence still required and the person who will return with it. Bring specialist judgement into the decision where required while retaining business ownership of the outcome.
Design for interruption and workload
Estimate review volume, response time and the operational impact when exceptions arrive together. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. A control that depends on an unavailable expert is not a functioning control. Record the choice, the evidence still required and the person who will return with it. Keep the mechanism proportionate: the purpose is better judgement and follow-through, not additional ceremony.
Create correction paths
Make it easy to amend the output, record the reason and identify patterns requiring prompt, data or process changes. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. Repeated human correction should improve the system rather than become permanent hidden labour. Record the choice, the evidence still required and the person who will return with it. Bring specialist judgement into the decision where required while retaining business ownership of the outcome.
Separate approval from accountability
Clarify who can release an item and who remains accountable for the design and performance of the workflow. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. Do not move organisational responsibility onto the most junior person clicking approve. Record the choice, the evidence still required and the person who will return with it. Keep the mechanism proportionate: the purpose is better judgement and follow-through, not additional ceremony.
Review the oversight model
Track error, intervention, false escalation, missed escalation, user effort and downstream outcomes over time. This matters because human in the loop AI agents decisions rarely fail through a lack of possible technology. They fail when the business problem, operating context and responsibility for the outcome remain implicit. Bring evidence from the people doing the work, the systems supporting it and the leaders accountable for the result. Test assumptions about time, behaviour, data quality and implementation effort before treating them as facts. Reduce or strengthen review only when evidence supports the change. Record the choice, the evidence still required and the person who will return with it. Bring specialist judgement into the decision where required while retaining business ownership of the outcome.
What to carry into the work
- The review model should respond to harm and reversibility, not apply one blanket rule.
- Review becomes theatre when a person sees only a polished output and no evidence.
- A control that depends on an unavailable expert is not a functioning control.
- Repeated human correction should improve the system rather than become permanent hidden labour.