LLM Routing vs Cascading: Which Model Handles a Query?

Routing chooses a model before it generates an answer. Cascading starts with one model and escalates after evaluating its answer. Use routing when the query supplies enough information to choose reliably. Consider cascading when checking a completed answer provides a better decision signal.

Both can reduce expensive-model calls. Their mistakes and latency costs differ.

When does the decision happen?

ApproachDecision signalMain failureCost to include
RoutingQuery features and predicted model capabilityA difficult query goes to an insufficient modelRouter and selected model
CascadingA generated answer and its evaluationA wrong answer passes the checkInitial model, evaluator, and escalated calls

A router might send a familiar FAQ question to a small model and a complex analysis to a stronger one. A cascade might accept a cheap extraction only after validating its fields against the source, then escalate failures.

Confidence stated by the answering model is not sufficient validation. Prefer a check tied to task correctness, such as verified source fields or tests with adequate coverage.

The RouteLLM paper’s Table 6 reports estimated cost divided by 3.66 on its MT-Bench setup at 95% of its GPT-4 quality score. That is about 72.7% lower cost under its historical model prices and evaluation setup. It is evidence for that experiment, not a current saving forecast.

FrugalGPT studies cascades across a model pool. Its best reported savings likewise depend on the tasks, model choices, and acceptance scoring.

Measure accepted mistakes separately

Evaluate easy and difficult cases, including queries where the cheap model produces a plausible wrong answer. Record the fraction handled cheaply, incorrect cheap answers accepted, escalations, total quality, and end-to-end latency.

For cascades, include both calls on escalated requests. For routing, include router latency and misrouted requests. Compare these totals with one-model baselines at the quality requirement you actually need.

The routing-and-cascading section of the LLM Engineering Guide illustrates both decision points.