04

Streaming

Set stream: true to receive server-sent events. Bytes pass straight through from the provider; usage arrives in the last chunk.

How it works

Chunks are standard OpenAI chat.completion.chunk objects. NXIO forces stream_options.include_usage so the final data event before [DONE] carries the usage object with cost. Billing happens after the stream closes, using that usage.

TypeScript
const stream = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [{ role: 'user', content: 'Write a haiku about servers.' }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
  if (chunk.usage) console.log('\n', chunk.usage); // final chunk carries usage + cost
}
raw SSE
data: {"id":"gen-…","choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}

data: {"id":"gen-…","choices":[{"index":0,"delta":{"content":"lo"},"finish_reason":"stop"}]}

data: {"id":"gen-…","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14,"prompt_tokens_details":{"cached_tokens":10,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000066,"is_byok":false}}

data: [DONE]

Errors during a stream

Once headers are sent the HTTP status is already 200. If the provider fails mid-stream, NXIO emits a final chunk with finish_reason "error" and a top-level error object, then [DONE]. Requests that fail before any bytes are sent return a normal HTTP error.

mid-stream error
data: {"id":null,"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"code":502,"message":"Upstream provider rejected the request","metadata":{"error_type":"server","provider_name":"upstream"}}}

data: [DONE]

If the provider omits usage

Rarely a stream ends without a usage chunk (connection cut, provider bug). NXIO then bills from a token estimate and marks the request usage_source = estimated in GET /generation. Estimates are deliberately conservative.