On CueCloud.
Connect to our endpoint. We run the model infrastructure; your team manages the application and decides which workloads can use the service.
Explore the hosted API ↗HOSTED INFERENCE / OPEN MODELS
Hosted inference for coding agents and applications. Connect through CueCode or an OpenAI-compatible API. We operate the infrastructure, with selected models priced 50% below comparable hosted rates.
Requests run on CueCloud infrastructure. Choose workloads approved for hosted processing.
Inference runs on infrastructure we operate.
CONNECT YOUR WORKFLOW
Request hosted access. Once enabled, get your API key from the console.
Use the base URL and the model ID provided for your enabled workspace. Keep your key out of source control.
Try a small, approved coding task before connecting a repository or agent workflow.
curl https://api.cuecloud.io/v1/chat/completions \
-H "Authorization: Bearer cue_••••••••" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_ENABLED_MODEL_ID",
"messages": [
{ "role": "user", "content": "fix the failing test" }
]
}'Example only. Replace the key and YOUR_ENABLED_MODEL_ID with your enabled workspace values.
HOSTED MODEL CATALOG
Selected open-weight models offered through our hosted service. Access depends on your enabled workspace; this catalog is not a live capacity monitor.
Hosted inference for your applications and coding workflows.
Your API model ID, request limits, and applicable rates are confirmed during onboarding.
HOSTED INFERENCE / LOWER COST
Inference for selected open-weight models at 50% less than comparable hosted rates from model labs and neocloud providers.
Enter a comparable hosted cost for a selected model. This illustration applies the 50% price difference; it does not quote a live provider rate or set billing terms.
$500 less for the same comparison period.
Match the model version, input/output mix, caching, and service terms. Eligibility and the comparison rate are confirmed with your team.
CHOOSE WHERE INFERENCE RUNS
Connect to our endpoint. We run the model infrastructure; your team manages the application and decides which workloads can use the service.
Explore the hosted API ↗Vault supplies local inference. Blitz brings the agent workspace, CueCode IDE, execution harness, and telemetry governance into your private environment.
Explore Vault and Blitz ↗YOUR NEXT STEP
Request hosted access and tell us about your workload. We’ll follow up on availability and onboarding. Creating a console account is a separate step and does not itself grant inference access.
Create a console account
Already have an account? Sign in ↗