How do you reduce LLM costs without losing answer quality?

Measure cost per successful task, then test output limits, caching, batching, or cheaper models on a held-out evaluation set. A smaller token bill is useful only if the system still meets its quality and latency requirements. Start with the repeated work or expensive request groups visible in your usage data.

Establish the actual cost

For an API with per-million-token prices, a simplified uncached request cost is:

cost=input tokens×input price+output tokens×output price106\text{cost} = \frac{\text{input tokens} \times \text{input price} + \text{output tokens} \times \text{output price}}{10^6}

Use current rates, such as those in the OpenAI pricing documentation. Account for cached-input prices, tool charges, and retries separately. Compare the sum across all calls needed to finish a task, including routing and validation.

For self-hosting, include GPU time, idle capacity, storage, networking, operations, and engineering. Dividing only GPU rental by generated tokens can hide substantial costs.

Match the change to eligible traffic

ChangeWhere it can helpWhat to verify
Shorter outputsResponses with unnecessary detailCompleteness and truncation errors
Prefix cachingRepeated compatible prompt prefixesHit rate, provider terms, and retained memory
Batch processingWork that can waitCompletion window and failure handling
Model routingRequests a cheaper model can answerMisrouted requests, router cost, and end-to-end quality
QuantizationSupported self-hosted model/runtime pairsQuality, throughput, memory, and usable batch size

Do not remove retrieved evidence merely to shorten a prompt. Check whether the missing context increases wrong answers or retries.

Compare complete tasks

Suppose a hypothetical workload sends 70% of requests to a cheaper model and 30% to a stronger model. Its model-call cost is 0.7 × cheap cost + 0.3 × strong cost, plus routing, retries, and other calls. Those are traffic assumptions, not a predicted saving.

Run the same evaluation set before and after the change. Report quality errors, latency, completion rate, and total cost. RouteLLM illustrates quality-constrained routing in specific experiments; its historical cost ratio does not predict your invoice.

Read cost optimization in the LLM Engineering Guide for the related serving techniques.