| Wire format |
|---|
| Content-Type | The response header that says 'this is SSE'. | text/event-stream |
| data: line | Each event carries its payload after 'data: '. | data: {"choices":[…]} |
| [DONE] | Sentinel most chat APIs send to close the stream. | data: [DONE] |
| event: field | Optional event type; LLM APIs mostly use default messages. | event: ping |
| id: / retry: | Reconnect hints — browsers auto-reconnect SSE; set retry in ms. | retry: 3000 |
| \n\n | Events are separated by a blank line — miss it and you merge events. | data: {…}\n\n |
| The chunk every provider converges on |
|---|
| delta.content | The text increment — concatenate them for the message. | delta: { content: "Hel" } |
| finish_reason | Null while streaming; 'stop', 'length', 'tool_calls' at the end. | finish_reason: 'stop' |
| usage | Token counts arrive in the final chunk (when requested). | usage: { prompt_tokens: 12 } |
| tool_call deltas | Function calls stream in id/name/arguments fragments — buffer and join. | delta: { tool_calls: […] } |
| role (first chunk) | The first delta carries the role; later ones carry content. | delta: { role: 'assistant' } |
| Client patterns |
|---|
| EventSource | Native but GET-only and header-less — LLM APIs usually can't use it. | new EventSource(url) |
| fetch + ReadableStream | The real pattern: POST with auth, parse the body stream. | for await (chunk of resp.body) |
| buffer split | Split on blank lines; keep the tail — a chunk can end mid-event. | buffer += chunk; split('\n\n') |
| abort | Stop generation with AbortController; servers bill only what ran. | controller.abort() |