How should LLM APIs limit requests, tokens, and concurrency?

An LLM API needs separate limits for request count, token volume, and concurrent work. Requests per minute alone cannot protect a service when prompt lengths and generated outputs vary. Add per-request input and output caps, then test admission limits against the latency target.

Why request counts are insufficient

A 10-token prompt and a 100,000-token prompt differ in input length by a factor of 10,000. Their total costs also depend on output length, model, and caching. A limit that treats them equally can admit more work than the service can process.

Use several controls together:

LimitWhat it controls
Requests per minuteCall volume, including many small requests
Input and output tokensVariable processing volume and spend
Concurrent requestsActive work that consumes decode capacity and KV memory
Tokens per requestIndividual prompts or outputs that exceed tested bounds

Apply tenant limits before a shared service-wide limit so one tenant cannot consume all available capacity.

How token reservation works

One application policy reserves estimated input tokens plus the maximum allowed output at admission. A request with 2,000 input tokens and an output cap of 1,000 reserves 3,000 tokens. If it produces 200 output tokens, actual usage is 2,200; the application refunds 800.

Keep that example separate from provider accounting. OpenAI documents its own request and token quota estimates. Anthropic documents separate input and output token limits. A local reservation does not predict which provider quota a request consumes. Reconcile interrupted requests when usage becomes available, and cap reservations that cannot be reconciled.

What to check under load

Mix short and long prompts, short and long outputs, and simultaneous tenants. Record queue time, rejected requests, TTFT, TPOT, and KV-cache use. Queue only within a bounded wait; otherwise reject with a clear retry policy. A tokens-per-minute budget can still admit a burst that exceeds concurrent memory capacity.

For the wider design, read rate limiting in the LLM Engineering Guide.