Key idea: Prompt injection hides instructions inside content the model reads. Check what comes back from tools, not just what the user types, and limit what the agent is allowed to do.
- Ask for a summary with the defaults. The input guardrail is on, but the attack is hiding in an email, not in your message.
- Turn on Scan tool results and ask again. The poisoned email never reaches the model.
- Scanning off, Mark untrusted content on. The attack stops, but notice the trade-off: the agent no longer acts on any email's request by itself.
- Set Send email to Ask first. Even a fooled agent can't send anything without you.
- Send the last prompt, a direct attack. This is the case the input guardrail is built for.
Try a prompt