Key idea: An eval is a fixed test set plus graders, giving a score you can compare across changes.
Prompt under test
Graders and model
- Run v1, then v2. Watch the format pass rate jump when the prompt spells out the output.
- Run v2, then v3, and compare the angry ticket and the "ignore all instructions" ticket.
- Switch to the Smart model and rerun v1. Does a bigger model fix a vague prompt?
- Turn off the LLM judge. The run is faster and cheaper, but what does it miss?
- Edit the prompt yourself and try to get 8/8.