06
Usage & billing
Every request is itemised: uncached input, cached input, cache writes, output. You are charged exactly that, in USD, from prepaid credits.
Formula
cost = (prompt_tokens − cached_tokens − cache_write_tokens) × input_price + cached_tokens × cached_input_price + cache_write_tokens × cache_write_price + completion_tokens × output_price. Reasoning tokens are part of completion_tokens. Per-request models charge pricing.request per call instead.
Prices are USD per token from GET /models. Streaming and non-streaming cost the same.
Credits, holds and settlement
Your balance is a ledger: top-ups are positive entries, charges negative. Before a request is forwarded a hold is placed for its worst-case cost; when the response completes the hold is released and the actual cost charged. If a provider reports more tokens than estimated the balance can dip slightly negative; the next request is then rejected with 402 until you top up.
Auditing a request
GET /generation?id=<request id> returns the stored token counts, the cost, latency and whether usage came from the provider or an estimate. The numbers are the same ones shown in your dashboard.
curl "https://api.nxioai.com/api/v1/generation?id=gen-mued4okl-6a7778ef4d23" \ -H "Authorization: Bearer $NXIO_API_KEY"