Key idea: Letting an agent act on its own is fast until it refunds the wrong customer. Asking a person every time is safe but slow. Route only risky actions (money, anything irreversible) to a person, and keep the policy in the agent's prompt so less reaches them at all.
100
Harm vs. a person's time
Auto-approve allAsk above $Money or irreversibleAsk for everything
- Run Auto-approve all. How many bad actions went through? Look for the $5,000 sent to an unknown bank account.
- Run Ask for everything. Nothing bad gets through, but count the reviews: a person approved every routine reply.
- Run Ask above $ at $100. Which bad actions still slipped through, because they didn't involve money?
- Run Money or irreversible. Where does its dot land between the two extremes?
- Turn on Give the agent the full policy and rerun Auto-approve all. How much does the policy prevent on its own?
- With the policy on, switch the agent to Smart. Is a bigger model automatically safer?