Model Routing

Key idea: Most requests don't need your biggest model. Route easy ones to a cheap model and hard ones to a strong one, and you keep most of the quality for a fraction of the cost. The catch: hard requests that look easy.

0.95

Accuracy vs. cost

Always smallAlways largeRouterCascade
0%25%50%75%100%$0$0.04$0.08$0.11Cost per 1,000 requestsRun a strategy to plot it here
  1. Run Always small, then Always large. Note the gap in accuracy and in cost per 1,000 requests.
  2. Run Router, then Cascade. Where do their dots land between the two?
  3. Lower the threshold to 0.5 and rerun Router. Find the short question it now calls easy: short questions can hide hard math.
  4. Look at Cascade's results: Small says it's sure even when it's wrong, so little escalates. Push the threshold to 1.0 and rerun. Self-reported confidence is a weak signal.
  5. Try a few thresholds and watch each strategy's dot slide between cheap and accurate.
escalateRequests12 support questionsRouteroffSmall modelcheap, fastgpt-4.1-nanoLarge modeloffGraderchecks the answerScorecardaccuracy and cost
Press Run to watch it run

Results

Run a strategy to see how each request was handled.

Behind the scenes

Run the requests and every routing decision will show up here.