How should LLM APIs limit requests, tokens, and concurrency?
An LLM API needs separate limits for request count, token volume, and concurrent work. Requests per minute alone cannot protect a service when prompt lengths and generated outputs vary. Add per-request input and output caps, then test admission limits against the latency target.
Why request counts are insufficient
A 10-token prompt and a 100,000-token prompt differ in input length by a factor of 10,000. Their total costs also depend on output length, model, and caching. A limit that treats them equally can admit more work than the service can process.
Use several controls together:
| Limit | What it controls |
|---|---|
| Requests per minute | Call volume, including many small requests |
| Input and output tokens | Variable processing volume and spend |
| Concurrent requests | Active work that consumes decode capacity and KV memory |
| Tokens per request | Individual prompts or outputs that exceed tested bounds |
Apply tenant limits before a shared service-wide limit so one tenant cannot consume all available capacity.
How token reservation works
One application policy reserves estimated input tokens plus the maximum allowed output at admission. A request with 2,000 input tokens and an output cap of 1,000 reserves 3,000 tokens. If it produces 200 output tokens, actual usage is 2,200; the application refunds 800.
Keep that example separate from provider accounting. OpenAI documents its own request and token quota estimates. Anthropic documents separate input and output token limits. A local reservation does not predict which provider quota a request consumes. Reconcile interrupted requests when usage becomes available, and cap reservations that cannot be reconciled.
What to check under load
Mix short and long prompts, short and long outputs, and simultaneous tenants. Record queue time, rejected requests, TTFT, TPOT, and KV-cache use. Queue only within a bounded wait; otherwise reject with a clear retry policy. A tokens-per-minute budget can still admit a burst that exceeds concurrent memory capacity.
For the wider design, read rate limiting in the LLM Engineering Guide.