Connecting several language models to the same workflow does not automatically create economies of scale. The system must demonstrate that, before knowing the correct answer, it can assign suitable requests to a lower-cost model, escalate requests that need greater capability, and retain a clear exit when that judgment fails. Enterprises should procure a verifiable allocation mechanism, not merely an interface that switches models.
Establish model eligibility before choosing a model
Consider an enterprise knowledge assistant handling terminology explanations, specification comparisons and after-sales recommendations. The first task may have stable answers; the others may involve unit conversions, conflicting document versions or authority to act. If allocation relies only on prompt length, a short question such as whether equipment may run above its rated load can be classified as simple. Textual complexity, business consequences and available evidence are distinct dimensions.
Routing should therefore begin with deterministic eligibility rules: may the model process this data, does it support the required context, language and tools, and does it meet location and authorization requirements? Only eligible candidates should enter the cost-quality comparison. Tasks involving irreversible actions should retain business approval or human review. Assigning a request to a more capable model does not supply authorization.
This article uses lower-cost model and stronger model as candidate roles, without treating parameter count as a capability guarantee. A specialized model may perform better on a defined task. The enterprise should freeze the task boundary and acceptance criteria, then use its own samples to establish which model is suitable for each request class. The recommendations are engineering analysis informed by public research. Every numerical example is hypothetical, not an IDENIFE measurement or a supplier quotation.
Pre-generation routing and post-generation cascades incur different costs
Pre-generation routing observes the input and chooses one model to generate a response. A central RouteLLM design predicts the probability that a stronger model will win a preference comparison, then uses a threshold to allocate requests.[1] In an enterprise application, this score expresses a relative selection preference, not answer correctness. Two wrong answers can still differ in perceived quality.
A post-generation cascade first obtains an answer from a lower-cost model. It delivers that answer if checks pass, or calls another model otherwise. FrugalGPT studies model-calling strategies that incorporate estimates of response quality.[2] A cascade can use additional signals from the answer, format checks or evidence verification. However, an escalated request has already incurred the first generation and checking costs, as well as sequential waiting time.
The choice should follow the available signals. When task classes and input features are sufficiently stable, pre-generation routing is a reasonable mechanism to evaluate first. When a problem becomes visible only after an answer exists, a cascade may be appropriate. Deterministic checks help with required output structure, but valid formatting does not establish factual correctness. Both mechanisms need an explicit option to withhold automatic delivery.
A lower escalation rate does not establish better routing
At least three measurements are needed: the proportion of requests covered by the lower-cost path, the error rate among answers delivered through that path, and cases that should have escalated but did not. The first two borrow the risk-coverage perspective from selective prediction: accepting more samples must be assessed alongside risk within the accepted set.[3] Statistical guarantees in that research depend on its task, method and sampling assumptions; they cannot be transferred directly into correctness guarantees for open-ended generation.
The third measurement requires paired evaluation. Suppose an independent sample contains 200 requests where the lower-cost model fails acceptance and the stronger model passes, but the router sends 40 of these to the lower-cost model. The missed-escalation rate within this recoverable group is 20%. This is a constructed example, not an overall error rate. Requests on which both models fail must be tracked separately, because escalation cannot be assumed to resolve them.
Business consequences also need separate treatment. An error in an editable summary and an incorrect equipment-handling recommendation should not cancel each other out in an average score. Acceptance criteria should define automatic-delivery conditions for consequential tasks, with additional review near decision thresholds and outside the evaluated distribution. When model judges supply labels, methods for calibrating false approvals can help assess the judges themselves. Agreement between models is not independent evidence.
Use a reproducible cost ledger to establish actual savings
For a two-model pre-generation router, a simplified average calling cost is C = C_r + (1 − p)C_s + pC_h. C_r is routing cost per request; C_s and C_h are the respective generation-path costs; p is the proportion sent to the stronger model. This approximation requires comparable requests when estimating costs. A production ledger must account separately for each request's input, output, cache hits and retries, rather than assuming equal token lengths across the two paths.
Consider a hypothetical example concerned only with calling fees. Assume C_r = CNY 0.002, C_s = CNY 0.012, C_h = CNY 0.120, and p = 30%. Then C = CNY 0.0464. At 100,000 monthly requests, calling fees total CNY 4,640, versus CNY 12,000 when every request uses the stronger model. The difference is CNY 7,360, or approximately 61.3%. These prices are arithmetic assumptions and do not describe any actual model quotation.
Now assume a cascade incurs C_e = CNY 0.003 for each check, sends every request first to the lower-cost model and escalates 30% once. Its average calling cost is C_s + C_e + pC_h = CNY 0.051, or CNY 5,100 per month at the same volume. Equal escalation rates do not imply equal quality. These amounts alone cannot determine the better design; final acceptance rates and latency after escalation must also be compared.
If the pre-generation routing system additionally needs CNY 6,000 per month in fixed evaluation and maintenance spending, its CNY 7,360 calling savings leave only CNY 1,360 before incident handling and other costs. The traffic needed to cover that fixed amount is approximately 6,000 ÷ (0.120 − 0.0464), rounded up to 81,522 requests per month. This threshold applies only to the stated assumptions. Falling traffic, longer outputs or more escalation should trigger a new calculation.
A cheaper path can improve the average while extending tail latency
Another hypothetical example illustrates the path structure. Assume 20 milliseconds for routing, 500 milliseconds for lower-cost generation, 2,000 milliseconds for stronger-model generation and 150 milliseconds for cascade checking, temporarily excluding queues and network delays. The pre-generation paths take 520 and 2,020 milliseconds. The cascade takes 650 milliseconds without escalation and 2,650 milliseconds with escalation. These figures illustrate sequential overhead; they are not production performance predictions.
A real system should measure time to first useful output, time to a complete answer and time to a final actionable business result separately. In particular, weighting the two models' individual P95 latencies by traffic share does not produce the combined P95. End-to-end percentiles should be calculated from actual request traces, while preserving a separate distribution for the escalation path. Difficult requests may also be longer and require more review, linking difficulty with higher latency.
Peak traffic and failures also matter. When the stronger model is rate-limited, will requests queue, be deferred or move to a human channel? Quietly downgrading consequential requests to maintain responsiveness invalidates the original quality acceptance. For self-hosted services, service-capacity validation should include escalation traffic converging on one backend, rather than testing only the average request mix.
Evaluation must reveal meaningful differences, not just an aggregate score
Establish three comparable baselines: every request uses the lower-cost model, every request uses the stronger model, and the proposed routing policy. Use identical business samples and acceptance rules, recording quality, complete cost, latency and undelivered requests separately. Preserve paired model outputs for evaluating candidate routers, so that the value of each switch can be assessed. Observing only the selected model's production answer does not establish whether the alternative would have been better.
Slice the sample at least by task type, language, input length, evidence completeness, business consequence and time period. Separate the data used to train the router, select thresholds and perform final acceptance. Near-duplicate rewrites of the same customer event should not leak between sets. Rare but consequential requests can receive additional evaluation samples, but aggregate costs and coverage should be reweighted to reflect actual traffic.
Begin with shadow operation: record proposed paths and reasons without changing what users receive. Expand traffic gradually within low-consequence slices supported by evaluation. Audits must sample automatically accepted answers, not only escalations or complaints; otherwise, false approvals can remain invisible. When evidence is sparse, report the sample count and uncertainty. No errors observed so far does not mean zero risk.
Manage thresholds, models and traffic changes together
The RouteLLM authors' repository recommends calibrating thresholds on data resembling incoming requests, and notes that actual model-call proportions can differ from those in calibration data.[4] An enterprise cannot treat 30% escalation at launch as a permanent property. New product lines, language shifts, prompt-template changes and model-service revisions can all alter relative capability and cost.
Each delivered result should be traceable to model identifiers, routing-policy and threshold versions, request slice, selection reason, checking results and the final path. Logging should follow business data-minimization requirements. Where needed, retain redacted features or controlled references rather than indiscriminately storing sensitive bodies. Configuration rollback should restore a mutually compatible set of versions.
An operational dashboard should show cost changes, escalation rates, false approvals by slice, tail latency and abstentions or human handoffs together. If any deteriorates, first distinguish changes in traffic, models and routing instead of automatically adjusting thresholds to chase cheaper calls. A new slice with insufficient evidence may require pausing automatic delivery and returning to a validated path. That fallback should be rehearsed before launch.
Procure decision evidence and agree on exit conditions
A demonstration that switches models establishes only that the connection layer works. An acceptance package should also include frozen model and policy versions, independent evaluation slices and sample counts, path-specific costs and end-to-end latency, missed escalations and joint failures, plus records of failure handling. Without exportable evidence, an enterprise will struggle to explain changing costs or migrate to a different routing service.
Define continuation criteria in advance: evidence supports the quality floor for target slices, net savings remain after human and operating costs, peak capacity meets business deadlines, and a named owner can execute rollback. Exit conditions include inadequate evidence for critical slices, persistent missed escalations, operating costs consuming the savings, or an inability to reproduce acceptance after a supplier change. Numerical limits must follow business consequences and sample evidence; there is no universal best escalation percentage.
The most useful result of multi-model routing is a traceable, correctable business decision about which requests merit more inference resources. A model portfolio can become part of a reliable production system only when the lower-cost path's scope, error-detection mechanism and escalation consequences can all be explained.
References
- [1] Ong et al. — RouteLLM: Learning to Route LLMs with Preference Data (v4, 2025)
- [2] Chen, Zaharia and Zou — FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (TMLR, 2024)
- [3] Geifman and El-Yaniv — Selective Classification for Deep Neural Networks (NeurIPS, 2017)
- [4] lm-sys — RouteLLM Authors' Repository: Threshold Calibration
