How do you reduce LLM costs without losing answer quality?
Measure cost per successful task, then test output limits, caching, batching, or cheaper models on a held-out evaluation set. A smaller token bill is useful only if the system still meets its quality and latency requirements. Start with the repeated work or expensive request groups visible in your usage data.
Establish the actual cost
For an API with per-million-token prices, a simplified uncached request cost is:
Use current rates, such as those in the OpenAI pricing documentation. Account for cached-input prices, tool charges, and retries separately. Compare the sum across all calls needed to finish a task, including routing and validation.
For self-hosting, include GPU time, idle capacity, storage, networking, operations, and engineering. Dividing only GPU rental by generated tokens can hide substantial costs.
Match the change to eligible traffic
| Change | Where it can help | What to verify |
|---|---|---|
| Shorter outputs | Responses with unnecessary detail | Completeness and truncation errors |
| Prefix caching | Repeated compatible prompt prefixes | Hit rate, provider terms, and retained memory |
| Batch processing | Work that can wait | Completion window and failure handling |
| Model routing | Requests a cheaper model can answer | Misrouted requests, router cost, and end-to-end quality |
| Quantization | Supported self-hosted model/runtime pairs | Quality, throughput, memory, and usable batch size |
Do not remove retrieved evidence merely to shorten a prompt. Check whether the missing context increases wrong answers or retries.
Compare complete tasks
Suppose a hypothetical workload sends 70% of requests to a cheaper model and 30% to a stronger model. Its model-call cost is 0.7 × cheap cost + 0.3 × strong cost, plus routing, retries, and other calls. Those are traffic assumptions, not a predicted saving.
Run the same evaluation set before and after the change. Report quality errors, latency, completion rate, and total cost. RouteLLM illustrates quality-constrained routing in specific experiments; its historical cost ratio does not predict your invoice.
Read cost optimization in the LLM Engineering Guide for the related serving techniques.