How to stream LLM responses using SSE and usage chunks
To stream an LLM response, parse the server’s event protocol, assemble content deltas, and update the display as usable text arrives. Many APIs use Server-Sent Events, or SSE. One HTTP read can contain part of an event or several events. An event can contain multiple model tokens.
Client implementations need buffering and completion handling as well as a fast first update.
Parse events before model data
The SSE format uses lines such as data: and ends an event with a blank line. Network reads can split those lines or combine several events. Buffer incomplete text, recognize complete events, then parse the API-specific payload. Multiple data: lines belong to one event.
For OpenAI Chat Completions, content arrives as deltas that the client assembles into a message. Handle tool-call or structured deltas according to their fields rather than treating every event as prose. The [DONE] marker is an API convention, not a general SSE requirement.
A useful client processing order is:
- Decode incoming bytes without losing a split UTF-8 character.
- Buffer until a complete SSE event is available.
- Interpret that event using the selected API’s schema.
- Append content or structured fields to the assembled response.
- Update the display with an appropriate buffer boundary.
Separate rendering, completion, and usage
Render prose promptly when possible. Buffer incomplete Markdown constructs or code lines if immediate formatting causes repeated layout changes. Record request start and first generated content separately from connection-open time or a metadata-only event.
In OpenAI’s Chat API, stream_options: {"include_usage": true} requests a final usage chunk before [DONE]. That usage chunk has an empty choices array. Interrupted or cancelled streams may never deliver it, so missing usage cannot be interpreted as zero tokens.
OpenAI-compatible servers may implement different payloads or completion behavior. Test fragmented events, several events per read, empty deltas, cancellation, server errors, and missing final usage against the pinned implementation. Keep already received text and distinguish an incomplete response from successful completion.
Engineering Guide: streaming in practice connects client behavior with TTFT, TPOT, and prefill scheduling.