Human-in-the-Loop

Key idea: Letting an agent act on its own is fast until it refunds the wrong customer. Asking a person every time is safe but slow. Route only risky actions (money, anything irreversible) to a person, and keep the policy in the agent's prompt so less reaches them at all.

100

Harm vs. a person's time

Auto-approve allAsk above $Money or irreversibleAsk for everything
0 bad24605914Reviews a person had to doRun a policy to plot it here
  1. Run Auto-approve all. How many bad actions went through? Look for the $5,000 sent to an unknown bank account.
  2. Run Ask for everything. Nothing bad gets through, but count the reviews: a person approved every routine reply.
  3. Run Ask above $ at $100. Which bad actions still slipped through, because they didn't involve money?
  4. Run Money or irreversible. Where does its dot land between the two extremes?
  5. Turn on Give the agent the full policy and rerun Auto-approve all. How much does the policy prevent on its own?
  6. With the policy on, switch the agent to Smart. Is a bigger model automatically safer?
askyesnoautoTasks12 customer requestsAgentproposes one actionministral-8bPolicy gaterun, or ask first?Human reviewoffExecutethe action runsScorecardharm vs. reviews
Press Run to watch it run

Results

Run the tasks to see what the agent did with each one.

Behind the scenes

Run the tasks and every proposed action, and who approved it, will show up here.