productpricing.md

Pricing

$500 per pack per month = 1B tokens. Add a pack when you need more. Predictable OpEx vs unbounded OpenAI and Anthropic API spend on agent loops.

Pack pricing

$500/ pack / mo1B tokens this period

$0.50/M · $500 = 1B combined in+out. Add a pack when you need more.

pack map
workspace: acme
plan: $0.50/M · $500 = 1B combined in+out this period.
key:  cue_…

What a pack includes

APIs charge per token. We sell packs. $0.50/M · $500 is 1B tokens this month. Need more? Add a pack.

  • 1B tokens in and out, per pack, per month
  • Unused include does not roll over
  • Need more? Add a pack.

Uncached 70/30

Run selected frontier-class coding models for 73–78% less. Save more than 90% vs labeled proprietary frontier at the same uncached 70/30 mix.

Cue stays $0.50/M · $500 = 1B combined in+out. List rates fetched 2026-08-25.

RouteList in / outUncached 70/30vs Cue $0.50/M
Cue Cloud ($0.50/M · $500 = 1B combined in+out)$0.50 / $0.50$0.50Cue unit
GLM-5.2 (Z.ai list, uncached)$1.40 / $4.40$2.30Cue 78% lower
DeepSeek V4 Pro 0813 (Peak cache-miss list)$1.32 / $3.96$2.11Cue 76% lower
Kimi K2.7 Code (Kimi list, uncached)$0.95 / $4$1.86Cue 73% lower
Kimi K3 (Kimi K3 is on Cue. Model id ships with the public API list.)$3 / $15$6.60Cue 92% lower
DeepSeek V4 Flash 0731 (Peak cache-miss list)$0.44 / $1.32$0.70No Cue list win
Anthropic · Opus 5 (Different model, uncached)$5 / $25$11Cue 95% lower
Anthropic · Fable 5 (Different model, uncached)$10 / $50$22Cue 98% lower
OpenAI · GPT-5.6 Sol (Different model, uncached)$4 / $20$8.80Cue 94% lower
Anthropic · Opus 4.8 (Different model, uncached)$5 / $25$11Cue 95% lower
OpenAI · GPT-5.5 (Different model, uncached)$5 / $30$12.50Cue 96% lower
OpenAI · GPT-5.6 Terra (Different model, uncached)$2 / $12$5Cue 90% lower
GPT-5.6 Luna (Same $/M. Different model.)$0.20 / $1.20$0.50Same $/M. Different model.

Price vs intelligence

Same model. Same intelligence. 73–78% lower serving price.

Horizontal pairs are the same model, same AA Intelligence Index. Cue at $0.50/M. Hosted points use the uncached 70/30 list blends above. Diamonds are different models. Not a same-model pair. Y is AA Intelligence Index, BenchLM snapshot 2026-08-25. Fetched 2026-08-25.

Same model. Same intelligence. 73–78% lower serving price.GLM-5.2: hosted $2.30/M to Cue $0.50/M at AA Intelligence Index 52.6, Cue 78% lower. DeepSeek V4 Pro 0813: hosted $2.11/M to Cue $0.50/M at AA Intelligence Index 53.2, Cue 76% lower. Kimi K2.7 Code: hosted $1.86/M to Cue $0.50/M at AA Intelligence Index 43, Cue 73% lower. Kimi K3: hosted $6.60/M to Cue $0.50/M at AA Intelligence Index 59.7, Cue 92% lower. DeepSeek V4 Flash: hosted $0.70/M at AA Intelligence Index 51.8. Named open circle. Different model. No Cue list win. Opus 5: $11/M at AA Intelligence Index 63. Different model. Fable 5: $22/M at AA Intelligence Index 62.1. Different model. GPT-5.6 Sol: $8.80/M at AA Intelligence Index 58.9. Different model. Opus 4.8: $11/M at AA Intelligence Index 57.3. Different model. GPT-5.5: $12.50/M at AA Intelligence Index 56.3. Different model. GPT-5.6 Terra: $5/M at AA Intelligence Index 55. Different model. Luna is on the 70/30 table, not this plot. Fetched 2026-08-25.434751555963$0.50$1$2$5$10$20AA Intelligence IndexUncached 70/30 $/M (log)Cue 78% lowerGLM-5.2Cue 76% lowerV4 Pro 0813Cue 73% lowerKimi K2.7 CodeCue 92% lowerKimi K3DeepSeek V4 FlashOpus 5Fable 5GPT-5.6 SolOpus 4.8GPT-5.5GPT-5.6 TerraCue $0.50/M
  • Cue Cloud
  • First-party hosted list
  • Different model
  • DeepSeek V4 Flash is named on the plot. Peak cache-miss. No Cue list win.

Uncached list prices. 70% input / 30% output. Fetched 2026-08-25. X is first-party hosted list, not a catalog quote. Y is AA Intelligence Index, BenchLM snapshot 2026-08-25. AA Index is a general composite, not a coding-only bench. Do not mix MCP-Atlas onto this axis. Horizontal pairs are the same model. Diamonds are different models. Cue Cloud does not change model intelligence. Kimi K3 is on Cue. Model id ships with the public API list. K3 vs Moonshot list is Cue 92% lower. That is not the 73-78 band. Flash 0731 is a different model from Pro 0813. Flash X is peak cache-miss. Off-peak Flash list can sit under the Cue unit. No Cue list-win stamp on Flash. Luna is on the 70/30 table at the same $/M as Cue. Different model. Not plotted. Cue is $0.50/M · $500 = 1B combined in+out. AA: https://benchlm.ai/benchmarks/artificialanalysis. Rate sources: https://developers.openai.com/api/docs/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.z.ai/guides/overview/pricing · https://api-docs.deepseek.com/quick_start/pricing · https://platform.kimi.ai/ · https://openrouter.ai/models. Bench sources: https://benchlm.ai/benchmarks/artificialanalysis.

Uncached list prices. 70% input / 30% output. Fetched 2026-08-25. X is first-party hosted list, not a catalog quote. Y is AA Intelligence Index, BenchLM snapshot 2026-08-25. AA Index is a general composite, not a coding-only bench. Do not mix MCP-Atlas onto this axis. Horizontal pairs are the same model. Diamonds are different models. Not a same-model pair. Flash 0731 is a different model from Pro 0813. Flash X is peak cache-miss. Off-peak Flash list can sit under the Cue unit. No Cue list-win stamp on Flash. Luna is on the 70/30 table at the same $/M as Cue. Different model. Not plotted. Cue is $0.50/M · $500 = 1B combined in+out. AA: AA Intelligence Index (BenchLM). Rate sources: OpenAI pricing · Anthropic pricing · Z.ai pricing · DeepSeek pricing · Kimi platform · OpenRouter catalog.

Same tokens. Pack vs meter.

Same pile. Cue is $500 a pack. Labeled OpenAI and Anthropic bill by the million at the 70/30 mix above.

Inside one pack

250M tokens this month. Still one pack.

250Mtokens / mo
Cue Cloud$500
Anthropic$2,750
OpenAI$3,125
  • $2,250 savings vs Anthropic
  • $2,625 savings vs OpenAI

One pack · 1B

Full include: 1 pack = 1B tokens (in + out). $500 this period.

1Btokens / mo
Cue Cloud$500
Anthropic$11,000
OpenAI$12,500
  • $10,500 savings vs Anthropic
  • $12,000 savings vs OpenAI

Two packs · 2B

Add a pack. $500 × 2. Metered still scales with every turn.

2Btokens / mo
Cue Cloud$1,000
Anthropic$22,000
OpenAI$25,000
  • $21,000 savings vs Anthropic
  • $24,000 savings vs OpenAI

Uncached list prices. 70% input / 30% output. Fetched 2026-08-25. Sources: https://developers.openai.com/api/docs/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.z.ai/guides/overview/pricing · https://api-docs.deepseek.com/quick_start/pricing · https://platform.kimi.ai/ · https://openrouter.ai/models. Cue is $0.50/M · $500 = 1B combined in+out. Caching, batch, peak or off-peak, and harness overhead move real bills. Unused include does not roll over.

Why coding agents punish OpenAI and Anthropic bills

IDE agents re-read repos, call tools, retry, and hold long context. Every turn is more tokens, so OpenAI and Anthropic OpEx grows without a ceiling. On Cue Cloud it is $500 × packs for that include.

Past ~45M blended tokens/mo on Anthropic (sooner on OpenAI at $12.5/MTok), one pack wins on dollars alone, plus open weights on hardware we operate. See models and api.

Pack anatomy

Packs, not people

Price is $500 × packs. Buy the include you need.

One cue_… key

Mint a cue_… key in the console. Use it in CueCode, Cursor, or any OpenAI-compatible client at https://api.cuecloud.io/v1.

Team math you can explain

N packs × $500 = this period’s include. The org Admin picks how many packs.

1 pack · 1B$500/mo
5 packs · 5B$2,500/mo
20 packs · 20B$10,000/mo

Console for humans

Manage packs and keys in the Cue Cloud console at app.cuecloud.io.

FAQ

Is the include capped?
Yes. One pack is 1B tokens in plus out this month. Unused tokens do not roll over. Need more? Add a pack.
Does paying turn inference on?
No. Request access first. After we enable the workspace, a pack is what you spend.
Who pays? Each engineer?
No. The org Admin buys packs for the workspace. Teammates use that include. They do not buy their own.
Can we share one cue_… key?
No. Each person should create their own key. A shared key makes spend impossible to attribute.
How are the OpenAI and Anthropic numbers calculated?
Published uncached list prices, fetched 2026-08-25. Opus 4.8 is $5 in / $25 out. GPT-5.5 is $5 in / $30 out. We mix 70% input / 30% output, then multiply by the monthly token volume in each scenario. Same volume on Cue is $500 times packs. Those are list math, not invoices. The dated table and chart above have the mix, the other routes, and the source links.
Back to explorer
change theme