04
Streaming
Set stream: true to receive server-sent events. Bytes pass straight through from the provider; usage arrives in the last chunk.
How it works
Chunks are standard OpenAI chat.completion.chunk objects. NXIO forces stream_options.include_usage so the final data event before [DONE] carries the usage object with cost. Billing happens after the stream closes, using that usage.
const stream = await client.chat.completions.create({
model: 'anthropic/claude-sonnet-5',
messages: [{ role: 'user', content: 'Write a haiku about servers.' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
if (chunk.usage) console.log('\n', chunk.usage); // final chunk carries usage + cost
}data: {"id":"gen-…","choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}
data: {"id":"gen-…","choices":[{"index":0,"delta":{"content":"lo"},"finish_reason":"stop"}]}
data: {"id":"gen-…","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14,"prompt_tokens_details":{"cached_tokens":10,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000066,"is_byok":false}}
data: [DONE]Errors during a stream
Once headers are sent the HTTP status is already 200. If the provider fails mid-stream, NXIO emits a final chunk with finish_reason "error" and a top-level error object, then [DONE]. Requests that fail before any bytes are sent return a normal HTTP error.
data: {"id":null,"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"code":502,"message":"Upstream provider rejected the request","metadata":{"error_type":"server","provider_name":"upstream"}}}
data: [DONE]If the provider omits usage
Rarely a stream ends without a usage chunk (connection cut, provider bug). NXIO then bills from a token estimate and marks the request usage_source = estimated in GET /generation. Estimates are deliberately conservative.