Metered frontier
OpenAI / Anthropic · $/MTok
waiting to start…
then metered costs · every turn adds $
The inference cloud for coding agents
$0.50/M · $500 = 1B combined in+out. CueCode or API.
Product
The inference cloud for coding agents. $0.50/M · $500 = 1B combined in+out this period. Open a topic for the full breakdown.
OpenAI-compatible. Works anywhere — optimized for IDE coding agents.
POST /v1/chat/completionsmodel: "deepseek-v4-pro" | "deepseek-v4-flash-0731" | "kimi-k2.7-code" | "glm-5.2"$0.50/M · $500 = 1BWhy Cue Cloud
$0.50/M · $500 = 1B this period
Frontier APIs price coding agents like chat: every read, tool call, and retry lands on an OpenAI or Anthropic meter. Teams want agents on every desk; finance wants a ceiling. Cue Cloud is that ceiling: one pack, one price, CueCode or API.
OpenAI / Anthropic · $/MTok
waiting to start…
then metered costs · every turn adds $
live pack · $0.50/M
waiting to start…
upfront once · every turn +$0
Illustrative turn costs for one heavy agent session. Not a live invoice. Upfront cost first, then the same loop on both sides. Full monthly math on packs.
A pack is the pricing model
Open Hub weights. $0.50/M. We sell packs, not a markup on someone else's rate card. That is how a 1B include stays forecastable while agents stay concurrent and long-running.
Paths
Same capacity pool. Two ways in. CueCode is our IDE. The API drops into the IDE you already use.
Same Cue Cloud models and pack include either way.
Proof
Illustrative probe — shaped display paint so you can see the decode loop. Not a live upstream stream.
deepseek-v4-prosess_7tmsxsdisplay paint · shaped 16–18 /s · not billed rate
warming session…
Live display / Reply avg are demo paint speed, not billed upstream rate and not a throughput SLA. Shaped band about 16–18 /s until Cue Cloud is measured.
Teams
$0.50/M · $500 = 1B combined in+out this period. Keys live in the Cue Cloud console.