Skip to content

Streaming & SSE Explained

Cheatsheet of Server-Sent Events as LLM APIs use them: wire anatomy, the chunk shape every provider converges on, and client patterns for parsing deltas.

LLM streaming rides on Server-Sent Events: one long-lived HTTP response where the server pushes text deltas as they generate. Every provider's payload differs in name and none in shape.

Reference table · 15 entries
15 of 15 rows
Wire format
Content-TypeThe response header that says 'this is SSE'.text/event-stream
data: lineEach event carries its payload after 'data: '.data: {"choices":[…]}
[DONE]Sentinel most chat APIs send to close the stream.data: [DONE]
event: fieldOptional event type; LLM APIs mostly use default messages.event: ping
id: / retry:Reconnect hints — browsers auto-reconnect SSE; set retry in ms.retry: 3000
\n\nEvents are separated by a blank line — miss it and you merge events.data: {…}\n\n
The chunk every provider converges on
delta.contentThe text increment — concatenate them for the message.delta: { content: "Hel" }
finish_reasonNull while streaming; 'stop', 'length', 'tool_calls' at the end.finish_reason: 'stop'
usageToken counts arrive in the final chunk (when requested).usage: { prompt_tokens: 12 }
tool_call deltasFunction calls stream in id/name/arguments fragments — buffer and join.delta: { tool_calls: […] }
role (first chunk)The first delta carries the role; later ones carry content.delta: { role: 'assistant' }
Client patterns
EventSourceNative but GET-only and header-less — LLM APIs usually can't use it.new EventSource(url)
fetch + ReadableStreamThe real pattern: POST with auth, parse the body stream.for await (chunk of resp.body)
buffer splitSplit on blank lines; keep the tail — a chunk can end mid-event.buffer += chunk; split('\n\n')
abortStop generation with AbortController; servers bill only what ran.controller.abort()