Nexara docs

Compute, models and pricing — all in one place.

Compute is Nexara's currency. Every model in the catalog is billed from it at that model's own rate — this reference shows exactly what each model costs per million tokens, which models take image or video input, and how the allowance works.

Compute

Compute is the only currency inside Nexara. One unit of Compute is worth $0.000001 — so 100,000 Compute = $0.10 and 1,000,000 Compute = $1.00. You buy a balance, subscribe to a plan for a monthly allowance, or earn free Compute — and every model you talk to is billed from that same pool at the model's own rate.

1M Compute = $1.00

The fixed exchange rate — the dollar figure is only ever an internal basis; you always see Compute.

Per-model rates

Every model bills its actual provider rate — a cheap flash model costs a fraction of a frontier flagship.

12,500 Compute per image

AI image generation (GPT Image) costs a flat 12,500 Compute per image, regardless of size.

MiniMax M3 is free

One exception to the rule: MiniMax M3 is 100% free to use — it never touches your Compute.

How the allowance works

  • Daily allowance. Your plan refills a set amount of Compute every day (shown in Settings → Compute). Using a model draws from it at that model's rate.
  • Weekly hard cap. A weekly ceiling bounds how much of the allowance you can burn, so a runaway session can't exhaust a whole month in a day.
  • Top-ups. Beyond the allowance, you can buy more Compute — it sits on your balance and is spent the same way, at the same per-model rates.
  • Cached tokens are cheaper. When a model re-reads part of your conversation from its prompt cache, those input tokens bill at the model's cache-read rate — often a fraction of the full input rate.

Model pricing

All 94 models in the catalog, with their Compute cost per million tokens — input, output, and cached input. The dollar figure underneath each Compute price is the internal provider rate (1M Compute = $1). MiniMax M3 is completely free. Models marked Locked are temporarily unavailable at the provider — they stay listed for reference but can't be selected until it recovers.

About the model IDs: catalog IDs follow the router convention provider/name (e.g. anthropic/claude-sonnet-4.6, openai/gpt-5.6-luna) — the same format routers like OpenRouter use. The prefix identifies the provider family a model belongs to, not a claim that the provider publishes that exact retail name. Models are served through a mix of direct provider connections (Claude models through Anthropic, GPT models through OpenAI) and open model routers, depending on the model, and the chat always shows you which model actually served your reply.

Frontier

25 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
GPT-OSS-120B
openai/gpt-oss-120b
100K$0.10400K$0.40131K
MiniMax M3Free
minimax/minimax-m3:free
FreeFree1M
Grok 4.5
x-ai/grok-4.5
1.74M$1.745.27M$5.27500K
Grok 4.6
x-ai/grok-4.6
2M$2.006M$6.00500K
GPT-5.6 Luna
openai/gpt-5.6-luna
200K$0.201.05M$1.051M
GPT-5.6 Terra
openai/gpt-5.6-terra
2.05M$2.0511.75M$11.751M
GPT-5.6 Sol
openai/gpt-5.6-sol
5M$5.0030M$30.00500K$0.501M
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
3M$3.0015M$15.00300K$0.301M
Claude Sonnet 5
anthropic/claude-sonnet-5
2M$2.0010M$10.00200K$0.201M
Claude Opus 4.6
anthropic/claude-opus-4.6
5M$5.0025M$25.00500K$0.501M
Claude Opus 4.7
anthropic/claude-opus-4.7
5M$5.0025M$25.00500K$0.501M
Claude Opus 4.8
anthropic/claude-opus-4.8
5M$5.0025M$25.00500K$0.501M
Claude Opus 5
anthropic/claude-opus-5
5M$5.0025M$25.00500K$0.501M
Claude Fable 5
anthropic/claude-fable-5
10M$10.0050M$50.001M$1.001M
Kimi K3
moonshotai/kimi-k3
3M$3.0015M$15.00300K$0.30262K
Kimi K2.6Locked
moonshotai/kimi-k2.6
950K$0.954M$4.00262K
Qwen 3.8 Max
qwen/qwen3.8-max
2.55M$2.557.55M$7.551M
Qwen 3.7 Max
qwen/qwen3.7-max
2.55M$2.557.55M$7.551M
Qwen 3.6 Max (Preview)
qwen/qwen3.6-max-preview
1.35M$1.357.85M$7.85256K
Qwen 3.5 397B A17B
qwen/qwen3.5-397b-a17b
650K$0.653.65M$3.65256K
Qwen3 Max
qwen/qwen3-max
1.25M$1.256.05M$6.05256K
GLM 5
z-ai/glm-5
630K$0.631.95M$1.95200K
GLM 5.1
z-ai/glm-5.1
1.42M$1.424.42M$4.42200K
GLM 5.2
z-ai/glm-5.2
1.5M$1.504.52M$4.521M
GLM 5.3
z-ai/glm-5.3
1.4M$1.404.4M$4.40260K$0.261M

Reasoning

15 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
Step 3.7 Flash
stepfun/step-3.7-flash
100K$0.10300K$0.30256K
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
50K$0.05150K$0.15256K
DeepSeek V4 Flash 07.31
deepseek/deepseek-v4-flash-0731
40K$0.04130K$0.139K$0.011M
Xiaomi Mimo V2.5 ProLocked
xiaomi/mimo-v2.5-pro:free
430K$0.43870K$0.87128K
Xiaomi Mimo V2.5Locked
xiaomi/mimo-v2.5:free
430K$0.43870K$0.87128K
Kimi K2.5Locked
moonshotai/kimi-k2.5
570K$0.572.85M$2.85262K
Nemotron 3 Nano
nvidia/nemotron-3-nano
100K$0.10400K$0.401M
Nemotron 3 Super
nvidia/nemotron-3-super
250K$0.251M$1.001M
Nemotron 3 Ultra
nvidia/nemotron-3-ultra
500K$0.502M$2.001M
Qwen 3.6 27B
qwen/qwen3.6-27b
650K$0.653.65M$3.65256K
Qwen 3.6 35B A3B
qwen/qwen3.6-35b-a3b
298K$0.301.54M$1.53256K
GLM 4.7
z-ai/glm-4.7
640K$0.642.24M$2.24200K
DeepSeek V3.2
deepseek/deepseek-v3.2
280K$0.28420K$0.42131K
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
250K$0.251M$1.001M
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
600K$0.602.4M$2.401M

General

20 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
MiniMax M2
minimax/minimax-m2:free
300K$0.301.2M$1.20205K
MiniMax M2.1
minimax/minimax-m2.1:free
300K$0.301.2M$1.20205K
MiniMax M2.5
minimax/minimax-m2.5:free
300K$0.301.2M$1.20205K
MiniMax M2.7
minimax/minimax-m2.7:free
300K$0.301.2M$1.20205K
Ministral 14B
mistralai/ministral-14b
200K$0.20200K$0.20256K
Mistral Small 26.03
mistralai/mistral-small-2603
100K$0.10300K$0.30256K
Mistral Medium 3.5
mistralai/mistral-medium-3.5
350K$0.351.05M$1.05256K
Mistral Large 3
mistralai/mistral-large-2512
2M$2.006M$6.00256K
Ling 3.0 Flash
inclusion-ai/ling-3.0-flash
50K$0.05150K$0.15260K
Llama 3.3 70B Instruct
meta/llama-3.3-70b-instruct
150K$0.15600K$0.60131K
Llama 3.1 8B Instruct
meta/llama-3.1-8b-instruct
40K$0.04100K$0.1016K
Qwen 3.7 Plus
qwen/qwen3.7-plus
450K$0.451.65M$1.651M
Qwen 3.6 Plus
qwen/qwen3.6-plus
550K$0.553.05M$3.051M
Qwen 3.5 Plus
qwen/qwen3.5-plus
450K$0.452.45M$2.451M
Qwen Plus 07.28
qwen/qwen-plus-2025-07-28
450K$0.451.25M$1.251M
GLM 4.5
z-ai/glm-4.5
630K$0.632.23M$2.23131K
DeepSeek Chat V3.1
deepseek/deepseek-chat-v3.1
280K$0.28420K$0.42164K
Llama 3.3 Nemotron Super 49B
nvidia/llama-3.3-nemotron-super-49b
150K$0.15600K$0.60131K
SenseNova 6.7 Flash-Lite
sensenova/sensenova-6.7-flash-lite
20K$0.0280K$0.08262K
SenseNova 6.8 Flash-Lite
sensenova/sensenova-6.8-flash-lite
20K$0.0280K$0.08262K

Coding

8 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
Devstral Medium
mistralai/devstral-medium
500K$0.50650K$0.65256K
Codestral 25.08
mistralai/codestral-2508
300K$0.30900K$0.90256K
Laguna XS.2
poolside/laguna-xs.2
100K$0.10400K$0.40131K
Grok Build 0.1
x-ai/grok-build-0.1
1M$1.002M$2.00200K$0.20256K
GPT-5.3 Codex Spark
openai/gpt-5.3-codex-spark
150K$0.15600K$0.60128K
Kimi K2.7 CodeLocked
moonshotai/kimi-k2.7-code
700K$0.703.5M$3.50150K$0.15262K
Qwen3 Coder Plus
qwen/qwen3-coder-plus
1.05M$1.055.05M$5.051M
GLM 4.6
z-ai/glm-4.6
620K$0.622.22M$2.22200K

Speed

13 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
MiniMax M2.1 High-Speed
minimax/minimax-m2.1-highspeed:free
150K$0.15600K$0.60205K
MiniMax M2.5 High-Speed
minimax/minimax-m2.5-highspeed:free
150K$0.15600K$0.60205K
MiniMax M2.7 High-Speed
minimax/minimax-m2.7-highspeed:free
150K$0.15600K$0.60205K
Ministral 3B
mistralai/ministral-3b
50K$0.0550K$0.05128K
Ministral 8B
mistralai/ministral-8b
100K$0.10100K$0.10256K
Nemotron Nano 9B V2
nvidia/nemotron-nano-9b-v2
30K$0.03100K$0.10128K
Llama 3.2 1B Instruct
meta/llama-3.2-1b-instruct
20K$0.0250K$0.0516K
Llama 3.2 3B Instruct
meta/llama-3.2-3b-instruct
30K$0.0380K$0.0816K
Claude Haiku 4.5
anthropic/claude-haiku-4.5
1M$1.005M$5.00100K$0.10200K
Qwen 3.5 Flash
qwen/qwen3.5-flash
150K$0.15450K$0.451M
GLM 4.5 Air
z-ai/glm-4.5-air
220K$0.221.12M$1.12131K
GLM 5 Turbo
z-ai/glm-5-turbo
1.24M$1.244.04M$4.04200K
GLM 5.3 Flash
z-ai/glm-5.3-flash
150K$0.15500K$0.5030K$0.031M

Multimodal

13 models
ModelCompute / 1M inCompute / 1M outCompute / 1M cacheImageVideoContext
Gemini 3.6 FlashLocked
google/gemini-3.6-flash
1.5M$1.507.5M$7.501M
Gemini 3.7 FlashLocked
google/gemini-3.7-flash
380K$0.381.88M$1.8840K$0.041M
Gemini 3.5 FlashLocked
google/gemini-3.5-flash
1.5M$1.509M$9.001M
Gemini 3.1 ProLocked
google/gemini-3.1-pro
2M$2.0012M$12.001M
Gemini 3 FlashLocked
google/gemini-3-flash
500K$0.503M$3.001M
Gemini 2.5 FlashLocked
google/gemini-2.5-flash
300K$0.302.5M$2.501M
Gemini 2.5 ProLocked
google/gemini-2.5-pro
1.25M$1.2510M$10.001M
Qwen 3.5 Omni Plus
qwen/qwen3.5-omni-plus
1.45M$1.4511.05M$11.05128K
Qwen 3.5 Omni Flash
qwen/qwen3.5-omni-flash
450K$0.453.05M$3.05128K
Qwen3 VL Plus
qwen/qwen3-vl-plus
250K$0.251.65M$1.65256K
Qwen3 Omni Flash
qwen/qwen3-omni-flash
480K$0.481.71M$1.71128K
Ox Alpha
stealth/ox-alpha-free
500K$0.501.25M$1.251M
Nemotron 3 Nano Omni
nvidia/nemotron-3-nano-omni
50K$0.05150K$0.15256K

Prices are Compute per million tokens at each model's actual provider rate. The cache column only appears when the model's provider advertises a separate cache-read rate — otherwise cached input bills at the full input rate. Context is the model's maximum window in tokens.

Image & video input

Some models accept media directly in the conversation — not just text. Image input means you can attach a picture and have the model read it (screenshots, charts, documents). Video input goes further: the model understands frames of a video clip. Media input is billed at the same token rate as text — vision tokens count toward your Compute like any other input.

Image input25

GPT-5.6 SolClaude Haiku 4.5Claude Sonnet 4.6Claude Sonnet 5Claude Opus 4.6Claude Opus 4.7Claude Opus 4.8Claude Opus 5Claude Fable 5Gemini 3.6 FlashGemini 3.7 FlashGemini 3.5 FlashGemini 3.1 ProGemini 3 FlashGemini 2.5 FlashGemini 2.5 ProQwen 3.5 Omni PlusQwen 3.5 Omni FlashQwen3 VL PlusQwen3 Omni FlashOx AlphaGLM 5.3 FlashNemotron 3 Nano OmniSenseNova 6.7 Flash-LiteSenseNova 6.8 Flash-Lite

Video input10

Gemini 3.6 FlashGemini 3.7 FlashGemini 3.5 FlashGemini 3.1 ProGemini 3 FlashGemini 2.5 FlashGemini 2.5 ProQwen 3.5 Omni PlusQwen 3.5 Omni FlashQwen3 Omni Flash

Frequently asked

What is one unit of Compute worth?

$0.000001 — so 100,000 Compute is $0.10 and 1,000,000 Compute is $1.00. All prices on this page are Compute per million tokens at that fixed rate.

Why do different models cost different amounts of Compute?

Each model bills at its actual provider rate. A tiny flash model costs a fraction of a frontier flagship because that's what the provider charges Nexara — the Compute price is that rate scaled to the 1M-Compute-per-dollar ratio.

Is any model free?

MiniMax M3 is 100% free to use — it never draws from your Compute. Every other model in the catalog is billed from Compute.

What is the cache rate?

When a conversation repeats the same system prompt and earlier messages, the provider can serve that prefix from its prompt cache. Those cached input tokens bill at the model's cache-read rate (shown in the table when a model has one) — usually 5–10× cheaper than fresh input.

What happens when I run out of Compute?

You're blocked from new requests until the daily allowance refills or the weekly cap resets Monday, or until you top up your balance. The app always shows your remaining Compute and refuses a request before it starts if the balance can't cover it.

For developers

Building something with Nexara or NexaraClaw? The Agent Visualization SDK connects your own front-end — a game-like NPC view, a voice assistant UI — to your local agents over a small WebSocket protocol.