08
Limits
Three independent limits protect your account: balance, per-key spend caps and the in-flight budget.
Where a 402 comes from
| limit_source | Cause | Fix |
|---|---|---|
| nxio_credits | Balance cannot cover the worst-case cost of this request. | Top up, or lower max_tokens. |
| nxio_key_limit | This key has reached its USD limit for the current period. | Raise the limit or wait for the reset (UTC). |
| nxio_in_flight_budget | Too many concurrent requests are holding credits. | Wait for them to finish (Retry-After is set) or add credits. |
| nxio_key_rpm (429) | This key exceeded 60 requests in the current minute. | Back off for Retry-After seconds; spread traffic across keys or ask for a higher limit. |
Rate limits
Each API key may make 60 requests per minute across all authenticated endpoints (fixed one-minute windows, UTC). Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset (unix seconds). Exceeding the limit returns 429 with error_type rate_limit_exceeded, limit_source nxio_key_rpm and a Retry-After header. Higher limits are available on request.
Provider rate limits are passed through unchanged as 429 with error_type rate_limit_exceeded; the Retry-After header is always authoritative.
Payload and context
Request bodies over 10 MB are rejected with 413. Context limits are per model; see context_length in the catalog.