03
Chat completions
POST /chat/completions accepts the OpenAI request body and forwards it to the model provider.
Request
Required: model and messages (or prompt). Supported optional parameters are passed through unchanged: temperature, top_p, max_tokens, max_completion_tokens, stop, tools, tool_choice, response_format, seed, frequency_penalty, presence_penalty, reasoning_effort, stream.
OpenRouter-only routing fields (provider, route, models, transforms, plugins) are accepted and ignored so existing code keeps working.
curl https://api.nxioai.com/api/v1/chat/completions \
-H "Authorization: Bearer $NXIO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"messages": [{"role": "user", "content": "Say hello"}]
}'Model ids and aliases
Use the vendor/slug id from the catalog (anthropic/claude-sonnet-5). The bare upstream id (claude-sonnet-5) is accepted as an alias, case-insensitively. The response echoes the canonical NXIO id in model.
max_tokens and credit holds
Before forwarding, NXIO reserves the worst-case cost of the request: estimated prompt tokens plus max_tokens (or the model maximum when omitted). Setting a realistic max_tokens keeps the reservation small, which matters when your balance is low. The reservation is released and replaced by the actual cost as soon as the response completes.
Response
Standard OpenAI shape. id is the NXIO request id (gen-…); the same value is returned in the X-Request-Id header. usage is always present, including on streams.
"usage": {
"prompt_tokens": 1200,
"completion_tokens": 300,
"total_tokens": 1500,
"prompt_tokens_details": { "cached_tokens": 1000, "cache_write_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 },
"cost": 0.00546,
"is_byok": false
}