Key idea: Reasoning models think before they answer. A little thinking fixes multi-step problems; more costs time and money without helping, and easy questions never needed it.
Questions right vs. time
MinimalLowMediumHigh
- Run Minimal. Which hard questions does it miss, and how fast is it?
- Run Low. How many more does it get right, and what did that cost in time and tokens?
- Run High. Did all that extra thinking buy any accuracy over Low?
- Look at the easy questions across your runs. Did thinking ever help them?