06

Usage & billing

Every request is itemised: uncached input, cached input, cache writes, output. You are charged exactly that, in USD, from prepaid credits.

Formula

cost = (prompt_tokens − cached_tokens − cache_write_tokens) × input_price + cached_tokens × cached_input_price + cache_write_tokens × cache_write_price + completion_tokens × output_price. Reasoning tokens are part of completion_tokens. Per-request models charge pricing.request per call instead.

Prices are USD per token from GET /models. Streaming and non-streaming cost the same.

Credits, holds and settlement

Your balance is a ledger: top-ups are positive entries, charges negative. Before a request is forwarded a hold is placed for its worst-case cost; when the response completes the hold is released and the actual cost charged. If a provider reports more tokens than estimated the balance can dip slightly negative; the next request is then rejected with 402 until you top up.

Auditing a request

GET /generation?id=<request id> returns the stored token counts, the cost, latency and whether usage came from the provider or an estimate. The numbers are the same ones shown in your dashboard.

curl
curl "https://api.nxioai.com/api/v1/generation?id=gen-mued4okl-6a7778ef4d23" \
  -H "Authorization: Bearer $NXIO_API_KEY"