Human-in-the-loop: designing AI agents people trust
Autonomy isn't all-or-nothing. The art of a trustworthy agent is choosing exactly where a human should stay in the loop.
The debate about AI autonomy is usually framed as a dial from “fully manual” to “fully autonomous”. In practice, the best agents aren't set to a single number — they're autonomous in some places and supervised in others, by design.
Match oversight to stakes
Reading data, drafting a response, categorising a transaction — low stakes, easily reversible, fine to automate. Sending money, deleting records, emailing a customer something binding — high stakes, worth a human glance. The skill is drawing that line deliberately, per action.
Make approvals fast and informative
Human-in-the-loop fails when approvals are slow or blind. A good approval shows the human exactly what the agent intends to do, why, and what it's based on — so the decision takes seconds, not minutes. Oversight should feel like a quick confirmation, not a second job.
Earn autonomy with evidence
Start an agent supervised. As its track record on a given action accumulates — measured, not assumed — you can graduate that action to full autonomy with confidence. Trust is something an agent earns through evals, not something you grant on faith.
Autonomy where it doesn't matter, oversight where it does — that's what makes an agent both useful and safe.