Reasoning Models

Key idea: Reasoning models think before they answer. A little thinking fixes multi-step problems; more costs time and money without helping, and easy questions never needed it.

Questions right vs. time

MinimalLowMediumHigh
0/124/128/1212/120.0s1.5s3.1s4.6sSeconds per question (average)Run an effort level to plot it here
  1. Run Minimal. Which hard questions does it miss, and how fast is it?
  2. Run Low. How many more does it get right, and what did that cost in time and tokens?
  3. Run High. Did all that extra thinking buy any accuracy over Low?
  4. Look at the easy questions across your runs. Did thinking ever help them?
thinkplanQuestions8 hard, 4 easyModelreads the questiongpt-5-nanoThinkingoffAnswerfinal line onlyGraderright or wrong?Scorecardaccuracy, time, cost
Press Run to watch it run

Results

Run the questions to see each answer.

Behind the scenes

Run the questions and every answer, with how long the model took, will show up here.