03

Chat completions

POST /chat/completions accepts the OpenAI request body and forwards it to the model provider.

Request

Required: model and messages (or prompt). Supported optional parameters are passed through unchanged: temperature, top_p, max_tokens, max_completion_tokens, stop, tools, tool_choice, response_format, seed, frequency_penalty, presence_penalty, reasoning_effort, stream.

OpenRouter-only routing fields (provider, route, models, transforms, plugins) are accepted and ignored so existing code keeps working.

curl
curl https://api.nxioai.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NXIO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5",
    "messages": [{"role": "user", "content": "Say hello"}]
  }'

Model ids and aliases

Use the vendor/slug id from the catalog (anthropic/claude-sonnet-5). The bare upstream id (claude-sonnet-5) is accepted as an alias, case-insensitively. The response echoes the canonical NXIO id in model.

max_tokens and credit holds

Before forwarding, NXIO reserves the worst-case cost of the request: estimated prompt tokens plus max_tokens (or the model maximum when omitted). Setting a realistic max_tokens keeps the reservation small, which matters when your balance is low. The reservation is released and replaced by the actual cost as soon as the response completes.

Response

Standard OpenAI shape. id is the NXIO request id (gen-…); the same value is returned in the X-Request-Id header. usage is always present, including on streams.

usage
"usage": {
  "prompt_tokens": 1200,
  "completion_tokens": 300,
  "total_tokens": 1500,
  "prompt_tokens_details": { "cached_tokens": 1000, "cache_write_tokens": 0 },
  "completion_tokens_details": { "reasoning_tokens": 0 },
  "cost": 0.00546,
  "is_byok": false
}