Decision guide
Can a smaller, cheaper model replace the LLM we use today?
The short answer: only against a bar set in advance
A smaller model can replace your current one when it clears a quality threshold you set before measuring, on every slice of difficulty you care about, and does so at a lower cost per accepted result. Often the honest outcome is partial: the smaller model handles the easy majority and fails on a hard minority, which points to routing rather than replacement. Sometimes it fails where it matters and the correct decision is to keep the model you have. All three are defensible answers; what is not defensible is switching on the strength of an average score and a per-token price.
Seven questions that decide it
Answer these questions in order. Each one can end the investigation.
- What does “good enough” mean, in writing? State the acceptance rule for a single output (schema-valid and every field correct; answer supported by the cited passage; reviewer would send it unedited) and the threshold for the workload, per slice. If you cannot write this down, you are not ready to compare models.
- Where is the money actually going? If most spend comes from a small number of long-context or high-volume calls, the decision may only matter for that path. Measure the share of spend by request type before generalising.
- Does the candidate pass on the hard slices? Averages are dominated by easy inputs. If the candidate fails the threshold on the slices where errors are expensive, it cannot replace the current model outright.
- Did the candidate get a fair prompt? A prompt tuned on the large model is not neutral. Re-optimise it for the candidate before drawing conclusions.
- Is the cost per accepted result lower once failures are paid for? Include retries, fallbacks to the large model, and the human time spent on rejected outputs.
- Does latency still meet the requirement? Smaller models are often faster per token, but extra retries or a fallback hop can make p95 worse.
- Is routing worth its running cost? A router adds a component to operate, a classifier or rule set to maintain, and a second model to monitor for drift. It pays off only when the easy and hard slices are distinguishable cheaply at request time.
If question 3 fails on every hard slice and question 7 fails because the slices cannot be told apart in advance, keep the current model and revisit when prices or models change.
What is measured, inferred and unknown
| What we can say | |
|---|---|
| Measured by Skelf | Nothing on this question. Skelf has published no model-replacement or routing benchmark. |
| Design fact | Route-Switch is a self-hosted, OpenAI-compatible Go gateway that routes across providers by configured strategy, logs per-prompt success, cost and latency to DuckDB, and reruns MIPROv2 prompt optimisation on captured traces (claim). It publishes no latency, cost-saving or quality benchmark. |
| Well-established | Per-token prices differ widely across model sizes and providers and change over time; output quality is not uniform across input difficulty; prompts transfer imperfectly between models. |
| Inferred | Workloads with a long tail of hard inputs are the ones where averages mislead most. Structured extraction with a strict schema is often more forgiving of smaller models than open-ended reasoning, but this varies by schema and domain and must be checked. |
| Unknown | The size of the quality gap, and the saving, on your traffic. No public leaderboard answers this, because your inputs, prompts and acceptance rule are not in it. |
The experiment Skelf plans to publish takes two workloads with different failure modes — structured extraction and grounded question answering — and compares a large model, a smaller model with its own re-optimised prompt, and a routed configuration, against a quality threshold fixed before any run, reporting cost per accepted result per difficulty slice with the dates and prices used. Until it is published, we make no claim about the result.
For the operating trade-offs of routing, the Route-Switch post on the quality, cost and latency triangle sets out how to choose which one is the constraint and which is allowed to slip.
When a past comparison stops holding
- Prices move. A comparison is a dated snapshot. A price cut on the large model, or a rise on the small one, can reverse the conclusion without any change in quality.
- Model versions change underneath you. Hosted models are updated; a “same name” model can behave differently next month. Pin versions where the provider allows it and re-test when you cannot.
- Traffic drifts. If the mix of easy and hard requests shifts — a new customer segment, a new document type, a new language — the slice weights you measured no longer describe production.
- The acceptance rule changes. Tightening the definition of a good output (a new compliance field, a stricter citation requirement) invalidates the old threshold comparison.
- Volume is small. If the workload costs little, the engineering and monitoring effort of switching can exceed the saving. “Not worth it at this volume” is a valid result.
How to run the comparison on your own traffic
This method needs a log of real requests, a way to judge outputs, and a few days. It produces a result you can defend to finance and to the team that owns quality.
Build the evaluation set. Sample a few hundred real requests from recent logs, stratified so that rare but important request types are represented, not just frequent ones. Tag each with a difficulty slice using features you can see at request time: input length, number of entities or fields, presence of tables, language, whether the answer requires combining several passages. Hold the set fixed for the whole study.
Fix the threshold before running anything. For example: “accepted-output rate within an agreed margin of the current model on every slice, and no slice below an absolute floor”. Have the quality owner sign it. Write down the date, model identifiers and list prices you will use.
Establish the baseline. Run the current model with its current prompt, at production settings, at least three times over the set, so you know its own run-to-run variation. A candidate cannot be judged against a baseline whose noise you have not measured.
Give the candidate a fair chance. Run it first with the current prompt, then with a prompt re-optimised for it on a separate development split, never on the evaluation set. Version both prompts (Blogus and Promptel can do this) so the comparison can be rerun.
Score with paired comparisons. Because both models answer the same items, compare them item by item. A paired test for binary outcomes, such as McNemar’s test, is more sensitive than comparing two averages, and a bootstrap over items gives an honest interval for each slice.
Compute the money correctly. Cost per accepted result equals model spend, plus retry and fallback spend, plus the cost of human handling for rejected outputs, divided by the number of accepted outputs. Report p50 and p95 latency per accepted result, including retries.
Test routing only if the slices separate. If the candidate passes on some slices and fails on others, check whether a cheap request-time rule or classifier can tell them apart. Measure the router’s misroute rate: a hard request sent to the small model is a quality failure; an easy one sent to the large model is a cost failure.
Shadow before you switch. Run the chosen configuration on mirrored live traffic for a period, compare against the current model’s outputs, and only then move real traffic.
If the spend or the quality risk is large
When the spend is large, or a quality regression would reach customers, Skelf can run this as a fixed-scope benchmark on your traffic: agree the threshold and slices with your quality owner, re-optimise prompts fairly for each candidate, measure large, small and routed configurations with dated prices, and hand over the harness so you can repeat it after the next model release. The report states plainly if the answer is to keep the current model. See the LLM cost and quality benchmark.
Related service: LLM cost–quality benchmark · How we run an investigation · Evidence register
Questions buyers ask
- Why compare cost per accepted result instead of cost per call?
- Because a cheaper model that produces more unusable answers pays for them in retries, escalations and human review. Cost per accepted result divides everything you spent by the outputs you could actually use, which is the number the business pays.
- Is an average benchmark score enough to make the switch?
- No. A smaller model can match the average while failing badly on the rare, hard inputs that matter most. Compare per slice of difficulty and set a floor for each slice, not just for the overall score.
- Should we re-optimise the prompt for the smaller model?
- Yes, before you reject it. Prompts tuned for one model often under-serve another, so a fair comparison gives the candidate its own optimised prompt and versions both prompts so the result can be reproduced.
- Does Route-Switch show how much routing saves?
- No. Route-Switch is a self-hosted gateway that routes by configured strategy and logs per-prompt success, cost and latency to DuckDB, but it publishes no cost-saving, latency or quality benchmark. It gives you the instrumentation to measure the trade-off on your own traffic.
- How often should we re-check the decision?
- Whenever a model version, a price, your prompts or your traffic mix changes. Record the date, model identifier and price used in every comparison, because a conclusion that was right in one quarter can be wrong in the next.
Is this decision on your desk?
If this decision is consequential, we can measure it on your workload: llm cost–quality benchmark, scoped to your deadline and the evidence that would settle it.