Models.Priced per token.
Prices are USD per million tokens. Cached-input pricing applies where the provider supports prompt caching.
…
Prices are USD per million tokens. Cached-input pricing applies where the provider supports prompt caching.
…
Anthropic claude-opus-5-5: the Opus tier of Claude, Anthropic's most capable line for complex reasoning, agents and software engineering. Text and image input.
Anthropic claude-fable-5-1: the Fable tier above Opus in Anthropic's Claude 5 family, for the hardest reasoning, research and coding work. Text and image input.
Z.ai GLM-5.3: GLM flagship chat and reasoning model with strong coding and agent performance.
Anthropic claude-fable-5: the Fable tier above Opus in Anthropic's Claude 5 family, for the hardest reasoning, research and coding work. Text and image input.
OpenAI gpt-5.5: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
OpenAI gpt-5.4: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
Google gemini-3.1-flash-image-preview: Gemini image generation and editing. Priced per generated image.
Z.ai GLM-5: GLM flagship chat and reasoning model with strong coding and agent performance.
Anthropic claude-opus-4.6: the Opus tier of Claude, Anthropic's most capable line for complex reasoning, agents and software engineering. Text and image input.
ByteDance doubao-seedream-4-5: Seedream image generation. Priced per image.
Anthropic's most capable Claude 4.5 model for complex reasoning, software engineering and agentic workflows. Vision input, 200K context window, up to 64K output tokens. Released 2025-11.
Google gemini-3-pro-image-preview: Gemini image generation and editing. Priced per generated image.
Anthropic's fastest and most affordable Claude 4.5 model. Near-frontier coding and agentic performance at a fraction of the cost, with vision input, a 200K context window and up to 64K output tokens. Released 2025-10.
Anthropic's balanced Claude 4.5 model: strong at coding, long-running agents and computer use, with vision input, a 200K context window and up to 64K output tokens. Released 2025-09-29.
ByteDance doubao-seedance-2-0-260128: Seedance video generation. Priced per generation.
ByteDance doubao-seedance-2-5-260628: Seedance video generation. Priced per generation.
ByteDance doubao-seedream-5-0-260128: Seedream image generation. Priced per image.
ByteDance doubao-seedream-5-0-pro-260628: Seedream image generation. Priced per image.
Alibaba qwen3-max: Qwen general-purpose chat and coding model.
Z.ai GLM-5.3-Flash: fast, low-cost GLM tier for high-throughput chat and coding.
OpenAI gpt-6.1-sol: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
Anthropic claude-sonnet-5-5: the Sonnet tier of Claude, balancing capability, speed and cost for coding and everyday agent workloads. Text and image input.
OpenAI gpt-6-luna: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
OpenAI gpt-6-sol: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
DeepSeek deepseek-flash: the fast, low-cost DeepSeek tier with reasoning tokens reported separately.
OpenAI gpt-image-2.5-flare: image generation and editing model. Priced per generated image.
OpenAI gpt-image-2.5-sunburst: image generation and editing model. Priced per generated image.
OpenAI gpt-6-astra: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
Google gemini-3.8-flash: the fast, price-efficient Gemini Flash tier with long context and optional thinking. Text and image input.
Google gemini-3.7-flash: the fast, price-efficient Gemini Flash tier with long context and optional thinking. Text and image input.
Alibaba qwen3.8-max: Qwen general-purpose chat and coding model.
Anthropic claude-opus-5: the Opus tier of Claude, Anthropic's most capable line for complex reasoning, agents and software engineering. Text and image input.
Moonshot AI Kimi-K3: Kimi frontier chat model with long context and strong agentic tool use. Text and image input.
OpenAI gpt-5.6-sol: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
Google gemini-3.1-flash-lite-image: Gemini image generation and editing. Priced per generated image.
Z.ai GLM-5.2: GLM flagship chat and reasoning model with strong coding and agent performance.
Moonshot AI Kimi-K2.7-Code: Kimi variant tuned for software engineering and agentic coding.
MiniMax MiniMax-M3: MiniMax M-series model for agentic coding and long-horizon tasks.
Anthropic claude-opus-4-8: the Opus tier of Claude, Anthropic's most capable line for complex reasoning, agents and software engineering. Text and image input.
Google gemini-3.5-flash: the fast, price-efficient Gemini Flash tier with long context and optional thinking. Text and image input.
Alibaba qwen3.7-max: Qwen general-purpose chat and coding model.
Qwen (Alibaba) happyhorse-1.0-i2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) happyhorse-1.0-r2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) happyhorse-1.0-t2v: video generation / editing model. Priced per generation.
DeepSeek deepseek-v4-flash: the fast, low-cost DeepSeek tier with reasoning tokens reported separately.
DeepSeek deepseek-v4-pro: general-purpose DeepSeek chat model with strong coding and reasoning at low cost.
OpenAI gpt-image-2: image generation and editing model. Priced per generated image.
Moonshot AI Kimi-K2.6: Kimi frontier chat model with long context and strong agentic tool use. Text and image input.
Anthropic claude-opus-4-7: the Opus tier of Claude, Anthropic's most capable line for complex reasoning, agents and software engineering. Text and image input.
Z.ai GLM-5-Turbo: fast, low-cost GLM tier for high-throughput chat and coding.
Z.ai GLM-5V-Turbo: GLM vision-language model for image understanding and multimodal chat.
Z.ai GLM-5.1: GLM flagship chat and reasoning model with strong coding and agent performance.
Alibaba qwen3.6-plus: Qwen general-purpose chat and coding model.
Alibaba qwen3.5-plus: Qwen general-purpose chat and coding model.
MiniMax MiniMax-M2.7: MiniMax M-series model for agentic coding and long-horizon tasks.
Google gemini-3.1-flash-lite-preview: the lightest, lowest-latency Gemini tier for high-volume, cost-sensitive workloads. Text and image input.
Alibaba qwen3.5-flash: fast, low-cost Qwen tier for high-volume chat.
Anthropic claude-sonnet-4.6: the Sonnet tier of Claude, balancing capability, speed and cost for coding and everyday agent workloads. Text and image input.
Alibaba Qwen3.5-397B-A17B: Qwen general-purpose chat and coding model.
MiniMax MiniMax-M2.5: MiniMax M-series model for agentic coding and long-horizon tasks.
Moonshot AI Kimi-K2.5: Kimi frontier chat model with long context and strong agentic tool use. Text and image input.
Alibaba Qwen3-Max-Thinking: Qwen reasoning variant with extended thinking (reasoning tokens billed as output).
Google gemini-3-flash-preview: the fast, price-efficient Gemini Flash tier with long context and optional thinking. Text and image input.
DeepSeek DeepSeek-V3.2: general-purpose DeepSeek chat model with strong coding and reasoning at low cost.
Google gemini-3-pro-preview: the Gemini Pro tier for complex reasoning, coding and long-context analysis. Text and image input.
Google gemini-2.5-flash-image: Gemini image generation and editing. Priced per generated image.
DeepSeek R1: open-weight reasoning model that emits chain-of-thought (billed as reasoning tokens) before answering. Strong on math and code.
Anthropic claude-sonnet-5: the Sonnet tier of Claude, balancing capability, speed and cost for coding and everyday agent workloads. Text and image input.
ByteDance doubao-seed-1-8-251228: Doubao Seed chat model with text and image input.
ByteDance doubao-seed-2-1-pro-260628: Doubao Seed chat model with text and image input.
ByteDance doubao-seed-2-1-turbo-260628: Doubao Seed chat model with text and image input.
Google's fast, price-efficient Gemini 2.5 workhorse with a 1M-token context window, up to 65K output tokens and optional thinking. Accepts text and images. Released 2025-06.
Google gemini-3.1-flash-lite: the lightest, lowest-latency Gemini tier for high-volume, cost-sensitive workloads. Text and image input.
Google gemini-3.1-pro-preview: the Gemini Pro tier for complex reasoning, coding and long-context analysis. Text and image input.
Google omni-fast: low-latency multimodal model for audio and voice-style interactions.
Google omni-fast-v2v: low-latency multimodal model for audio and voice-style interactions.
OpenAI gpt-5.6-luna: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
OpenAI gpt-5.6-terra: a GPT-series chat and reasoning model available through the Chat Completions and Responses APIs. Text and image input.
Alibaba qwen-deep-research: agentic deep-research model that plans, browses and synthesises long reports.
Alibaba qwen-flash: fast, low-cost Qwen tier for high-volume chat.
Alibaba qwen3-vl-flash: Qwen vision-language model for image understanding and multimodal chat.
Alibaba qwen3-vl-plus: Qwen vision-language model for image understanding and multimodal chat.
Alibaba's Qwen text-embedding-v4: multilingual embeddings returning 1024-dimensional vectors (verified through NXIO on 2026-09-28). Billed on input tokens only.
Qwen (Alibaba) wan2.6-i2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan2.6-t2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan2.7-i2v: video generation / editing model. Priced per generation.
Alibaba wan2.7-image: Wan image generation. Priced per image.
Alibaba wan2.7-image-pro: Wan image generation. Priced per image.
Qwen (Alibaba) wan2.7-r2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan2.7-t2v: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan2.7-videoedit: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan3.0-video: video generation / editing model. Priced per generation.
Qwen (Alibaba) wan3.0-video-prime: video generation / editing model. Priced per generation.
Skywork (Kunlun) skyreels: text in, video out, priced per request, available through the NXIO gateway.
xAI grok-imagine-video-1.5-preview: video generation / editing model. Priced per generation.