Evaluating AI Changes

Key idea: An eval is a fixed test set plus graders, giving a score you can compare across changes.

Prompt under test

Graders and model

  1. Run v1, then v2. Watch the format pass rate jump when the prompt spells out the output.
  2. Run v2, then v3, and compare the angry ticket and the "ignore all instructions" ticket.
  3. Switch to the Smart model and rerun v1. Does a bigger model fix a vague prompt?
  4. Turn off the LLM judge. The run is faster and cheaper, but what does it miss?
  5. Edit the prompt yourself and try to get 8/8.
Test set8 labeled ticketsPromptversion under testModelanswers each caseministral-8bFormat checkvalid JSON?Exact matchright category?LLM judgereply quality 1-5ministral-8bScorecardpass rate
Press Run to watch it run

Results

Run the eval to see how each test case scores.

Behind the scenes

Run the eval and every step, case by case, will show up here.