Compute, models and pricing — all in one place.
Compute is Nexara's currency. Every model in the catalog is billed from it at that model's own rate — this reference shows exactly what each model costs per million tokens, which models take image or video input, and how the allowance works.
Compute
Compute is the only currency inside Nexara. One unit of Compute is worth $0.000001 — so 100,000 Compute = $0.10 and 1,000,000 Compute = $1.00. You buy a balance, subscribe to a plan for a monthly allowance, or earn free Compute — and every model you talk to is billed from that same pool at the model's own rate.
1M Compute = $1.00
The fixed exchange rate — the dollar figure is only ever an internal basis; you always see Compute.
Per-model rates
Every model bills its actual provider rate — a cheap flash model costs a fraction of a frontier flagship.
12,500 Compute per image
AI image generation (GPT Image) costs a flat 12,500 Compute per image, regardless of size.
MiniMax M3 is free
One exception to the rule: MiniMax M3 is 100% free to use — it never touches your Compute.
How the allowance works
- Daily allowance. Your plan refills a set amount of Compute every day (shown in Settings → Compute). Using a model draws from it at that model's rate.
- Weekly hard cap. A weekly ceiling bounds how much of the allowance you can burn, so a runaway session can't exhaust a whole month in a day.
- Top-ups. Beyond the allowance, you can buy more Compute — it sits on your balance and is spent the same way, at the same per-model rates.
- Cached tokens are cheaper. When a model re-reads part of your conversation from its prompt cache, those input tokens bill at the model's cache-read rate — often a fraction of the full input rate.
Model pricing
All 94 models in the catalog, with their Compute cost per million tokens — input, output, and cached input. The dollar figure underneath each Compute price is the internal provider rate (1M Compute = $1). MiniMax M3 is completely free. Models marked Locked are temporarily unavailable at the provider — they stay listed for reference but can't be selected until it recovers.
Frontier
25 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
GPT-OSS-120B openai/gpt-oss-120b | 100K$0.10 | 400K$0.40 | — | 131K | ||
MiniMax M3Free minimax/minimax-m3:free | Free | Free | — | 1M | ||
Grok 4.5 x-ai/grok-4.5 | 1.74M$1.74 | 5.27M$5.27 | — | 500K | ||
Grok 4.6 x-ai/grok-4.6 | 2M$2.00 | 6M$6.00 | — | 500K | ||
GPT-5.6 Luna openai/gpt-5.6-luna | 200K$0.20 | 1.05M$1.05 | — | 1M | ||
GPT-5.6 Terra openai/gpt-5.6-terra | 2.05M$2.05 | 11.75M$11.75 | — | 1M | ||
GPT-5.6 Sol openai/gpt-5.6-sol | 5M$5.00 | 30M$30.00 | 500K$0.50 | 1M | ||
Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 | 3M$3.00 | 15M$15.00 | 300K$0.30 | 1M | ||
Claude Sonnet 5 anthropic/claude-sonnet-5 | 2M$2.00 | 10M$10.00 | 200K$0.20 | 1M | ||
Claude Opus 4.6 anthropic/claude-opus-4.6 | 5M$5.00 | 25M$25.00 | 500K$0.50 | 1M | ||
Claude Opus 4.7 anthropic/claude-opus-4.7 | 5M$5.00 | 25M$25.00 | 500K$0.50 | 1M | ||
Claude Opus 4.8 anthropic/claude-opus-4.8 | 5M$5.00 | 25M$25.00 | 500K$0.50 | 1M | ||
Claude Opus 5 anthropic/claude-opus-5 | 5M$5.00 | 25M$25.00 | 500K$0.50 | 1M | ||
Claude Fable 5 anthropic/claude-fable-5 | 10M$10.00 | 50M$50.00 | 1M$1.00 | 1M | ||
Kimi K3 moonshotai/kimi-k3 | 3M$3.00 | 15M$15.00 | 300K$0.30 | 262K | ||
Kimi K2.6Locked moonshotai/kimi-k2.6 | 950K$0.95 | 4M$4.00 | — | 262K | ||
Qwen 3.8 Max qwen/qwen3.8-max | 2.55M$2.55 | 7.55M$7.55 | — | 1M | ||
Qwen 3.7 Max qwen/qwen3.7-max | 2.55M$2.55 | 7.55M$7.55 | — | 1M | ||
Qwen 3.6 Max (Preview) qwen/qwen3.6-max-preview | 1.35M$1.35 | 7.85M$7.85 | — | 256K | ||
Qwen 3.5 397B A17B qwen/qwen3.5-397b-a17b | 650K$0.65 | 3.65M$3.65 | — | 256K | ||
Qwen3 Max qwen/qwen3-max | 1.25M$1.25 | 6.05M$6.05 | — | 256K | ||
GLM 5 z-ai/glm-5 | 630K$0.63 | 1.95M$1.95 | — | 200K | ||
GLM 5.1 z-ai/glm-5.1 | 1.42M$1.42 | 4.42M$4.42 | — | 200K | ||
GLM 5.2 z-ai/glm-5.2 | 1.5M$1.50 | 4.52M$4.52 | — | 1M | ||
GLM 5.3 z-ai/glm-5.3 | 1.4M$1.40 | 4.4M$4.40 | 260K$0.26 | 1M |
Reasoning
15 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
Step 3.7 Flash stepfun/step-3.7-flash | 100K$0.10 | 300K$0.30 | — | 256K | ||
Nemotron 3 Nano 30B A3B nvidia/nemotron-3-nano-30b-a3b | 50K$0.05 | 150K$0.15 | — | 256K | ||
DeepSeek V4 Flash 07.31 deepseek/deepseek-v4-flash-0731 | 40K$0.04 | 130K$0.13 | 9K$0.01 | 1M | ||
Xiaomi Mimo V2.5 ProLocked xiaomi/mimo-v2.5-pro:free | 430K$0.43 | 870K$0.87 | — | 128K | ||
Xiaomi Mimo V2.5Locked xiaomi/mimo-v2.5:free | 430K$0.43 | 870K$0.87 | — | 128K | ||
Kimi K2.5Locked moonshotai/kimi-k2.5 | 570K$0.57 | 2.85M$2.85 | — | 262K | ||
Nemotron 3 Nano nvidia/nemotron-3-nano | 100K$0.10 | 400K$0.40 | — | 1M | ||
Nemotron 3 Super nvidia/nemotron-3-super | 250K$0.25 | 1M$1.00 | — | 1M | ||
Nemotron 3 Ultra nvidia/nemotron-3-ultra | 500K$0.50 | 2M$2.00 | — | 1M | ||
Qwen 3.6 27B qwen/qwen3.6-27b | 650K$0.65 | 3.65M$3.65 | — | 256K | ||
Qwen 3.6 35B A3B qwen/qwen3.6-35b-a3b | 298K$0.30 | 1.54M$1.53 | — | 256K | ||
GLM 4.7 z-ai/glm-4.7 | 640K$0.64 | 2.24M$2.24 | — | 200K | ||
DeepSeek V3.2 deepseek/deepseek-v3.2 | 280K$0.28 | 420K$0.42 | — | 131K | ||
DeepSeek V4 Flash deepseek/deepseek-v4-flash | 250K$0.25 | 1M$1.00 | — | 1M | ||
DeepSeek V4 Pro deepseek/deepseek-v4-pro | 600K$0.60 | 2.4M$2.40 | — | 1M |
General
20 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
MiniMax M2 minimax/minimax-m2:free | 300K$0.30 | 1.2M$1.20 | — | 205K | ||
MiniMax M2.1 minimax/minimax-m2.1:free | 300K$0.30 | 1.2M$1.20 | — | 205K | ||
MiniMax M2.5 minimax/minimax-m2.5:free | 300K$0.30 | 1.2M$1.20 | — | 205K | ||
MiniMax M2.7 minimax/minimax-m2.7:free | 300K$0.30 | 1.2M$1.20 | — | 205K | ||
Ministral 14B mistralai/ministral-14b | 200K$0.20 | 200K$0.20 | — | 256K | ||
Mistral Small 26.03 mistralai/mistral-small-2603 | 100K$0.10 | 300K$0.30 | — | 256K | ||
Mistral Medium 3.5 mistralai/mistral-medium-3.5 | 350K$0.35 | 1.05M$1.05 | — | 256K | ||
Mistral Large 3 mistralai/mistral-large-2512 | 2M$2.00 | 6M$6.00 | — | 256K | ||
Ling 3.0 Flash inclusion-ai/ling-3.0-flash | 50K$0.05 | 150K$0.15 | — | 260K | ||
Llama 3.3 70B Instruct meta/llama-3.3-70b-instruct | 150K$0.15 | 600K$0.60 | — | 131K | ||
Llama 3.1 8B Instruct meta/llama-3.1-8b-instruct | 40K$0.04 | 100K$0.10 | — | 16K | ||
Qwen 3.7 Plus qwen/qwen3.7-plus | 450K$0.45 | 1.65M$1.65 | — | 1M | ||
Qwen 3.6 Plus qwen/qwen3.6-plus | 550K$0.55 | 3.05M$3.05 | — | 1M | ||
Qwen 3.5 Plus qwen/qwen3.5-plus | 450K$0.45 | 2.45M$2.45 | — | 1M | ||
Qwen Plus 07.28 qwen/qwen-plus-2025-07-28 | 450K$0.45 | 1.25M$1.25 | — | 1M | ||
GLM 4.5 z-ai/glm-4.5 | 630K$0.63 | 2.23M$2.23 | — | 131K | ||
DeepSeek Chat V3.1 deepseek/deepseek-chat-v3.1 | 280K$0.28 | 420K$0.42 | — | 164K | ||
Llama 3.3 Nemotron Super 49B nvidia/llama-3.3-nemotron-super-49b | 150K$0.15 | 600K$0.60 | — | 131K | ||
SenseNova 6.7 Flash-Lite sensenova/sensenova-6.7-flash-lite | 20K$0.02 | 80K$0.08 | — | 262K | ||
SenseNova 6.8 Flash-Lite sensenova/sensenova-6.8-flash-lite | 20K$0.02 | 80K$0.08 | — | 262K |
Coding
8 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
Devstral Medium mistralai/devstral-medium | 500K$0.50 | 650K$0.65 | — | 256K | ||
Codestral 25.08 mistralai/codestral-2508 | 300K$0.30 | 900K$0.90 | — | 256K | ||
Laguna XS.2 poolside/laguna-xs.2 | 100K$0.10 | 400K$0.40 | — | 131K | ||
Grok Build 0.1 x-ai/grok-build-0.1 | 1M$1.00 | 2M$2.00 | 200K$0.20 | 256K | ||
GPT-5.3 Codex Spark openai/gpt-5.3-codex-spark | 150K$0.15 | 600K$0.60 | — | 128K | ||
Kimi K2.7 CodeLocked moonshotai/kimi-k2.7-code | 700K$0.70 | 3.5M$3.50 | 150K$0.15 | 262K | ||
Qwen3 Coder Plus qwen/qwen3-coder-plus | 1.05M$1.05 | 5.05M$5.05 | — | 1M | ||
GLM 4.6 z-ai/glm-4.6 | 620K$0.62 | 2.22M$2.22 | — | 200K |
Speed
13 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
MiniMax M2.1 High-Speed minimax/minimax-m2.1-highspeed:free | 150K$0.15 | 600K$0.60 | — | 205K | ||
MiniMax M2.5 High-Speed minimax/minimax-m2.5-highspeed:free | 150K$0.15 | 600K$0.60 | — | 205K | ||
MiniMax M2.7 High-Speed minimax/minimax-m2.7-highspeed:free | 150K$0.15 | 600K$0.60 | — | 205K | ||
Ministral 3B mistralai/ministral-3b | 50K$0.05 | 50K$0.05 | — | 128K | ||
Ministral 8B mistralai/ministral-8b | 100K$0.10 | 100K$0.10 | — | 256K | ||
Nemotron Nano 9B V2 nvidia/nemotron-nano-9b-v2 | 30K$0.03 | 100K$0.10 | — | 128K | ||
Llama 3.2 1B Instruct meta/llama-3.2-1b-instruct | 20K$0.02 | 50K$0.05 | — | 16K | ||
Llama 3.2 3B Instruct meta/llama-3.2-3b-instruct | 30K$0.03 | 80K$0.08 | — | 16K | ||
Claude Haiku 4.5 anthropic/claude-haiku-4.5 | 1M$1.00 | 5M$5.00 | 100K$0.10 | 200K | ||
Qwen 3.5 Flash qwen/qwen3.5-flash | 150K$0.15 | 450K$0.45 | — | 1M | ||
GLM 4.5 Air z-ai/glm-4.5-air | 220K$0.22 | 1.12M$1.12 | — | 131K | ||
GLM 5 Turbo z-ai/glm-5-turbo | 1.24M$1.24 | 4.04M$4.04 | — | 200K | ||
GLM 5.3 Flash z-ai/glm-5.3-flash | 150K$0.15 | 500K$0.50 | 30K$0.03 | 1M |
Multimodal
13 models| Model | Compute / 1M in | Compute / 1M out | Compute / 1M cache | Image | Video | Context |
|---|---|---|---|---|---|---|
Gemini 3.6 FlashLocked google/gemini-3.6-flash | 1.5M$1.50 | 7.5M$7.50 | — | 1M | ||
Gemini 3.7 FlashLocked google/gemini-3.7-flash | 380K$0.38 | 1.88M$1.88 | 40K$0.04 | 1M | ||
Gemini 3.5 FlashLocked google/gemini-3.5-flash | 1.5M$1.50 | 9M$9.00 | — | 1M | ||
Gemini 3.1 ProLocked google/gemini-3.1-pro | 2M$2.00 | 12M$12.00 | — | 1M | ||
Gemini 3 FlashLocked google/gemini-3-flash | 500K$0.50 | 3M$3.00 | — | 1M | ||
Gemini 2.5 FlashLocked google/gemini-2.5-flash | 300K$0.30 | 2.5M$2.50 | — | 1M | ||
Gemini 2.5 ProLocked google/gemini-2.5-pro | 1.25M$1.25 | 10M$10.00 | — | 1M | ||
Qwen 3.5 Omni Plus qwen/qwen3.5-omni-plus | 1.45M$1.45 | 11.05M$11.05 | — | 128K | ||
Qwen 3.5 Omni Flash qwen/qwen3.5-omni-flash | 450K$0.45 | 3.05M$3.05 | — | 128K | ||
Qwen3 VL Plus qwen/qwen3-vl-plus | 250K$0.25 | 1.65M$1.65 | — | 256K | ||
Qwen3 Omni Flash qwen/qwen3-omni-flash | 480K$0.48 | 1.71M$1.71 | — | 128K | ||
Ox Alpha stealth/ox-alpha-free | 500K$0.50 | 1.25M$1.25 | — | 1M | ||
Nemotron 3 Nano Omni nvidia/nemotron-3-nano-omni | 50K$0.05 | 150K$0.15 | — | 256K |
Prices are Compute per million tokens at each model's actual provider rate. The cache column only appears when the model's provider advertises a separate cache-read rate — otherwise cached input bills at the full input rate. Context is the model's maximum window in tokens.
Image & video input
Some models accept media directly in the conversation — not just text. Image input means you can attach a picture and have the model read it (screenshots, charts, documents). Video input goes further: the model understands frames of a video clip. Media input is billed at the same token rate as text — vision tokens count toward your Compute like any other input.
Image input25
Video input10
Frequently asked
What is one unit of Compute worth?
Why do different models cost different amounts of Compute?
Is any model free?
What is the cache rate?
What happens when I run out of Compute?
For developers
Building something with Nexara or NexaraClaw? The Agent Visualization SDK connects your own front-end — a game-like NPC view, a voice assistant UI — to your local agents over a small WebSocket protocol.
