HOSTED INFERENCE / OPEN MODELS

Open models.
Ready for your workload.

Hosted inference for coding agents and applications. Connect through CueCode or an OpenAI-compatible API. We operate the infrastructure, with selected models priced 50% below comparable hosted rates.

Connect your workflow ↓

Requests run on CueCloud infrastructure. Choose workloads approved for hosted processing.

REQUEST PATH / CONCEPT VIEW
YOUR ENVIRONMENT
Your applicationCueCode IDE
REQUEST ↓
CUECLOUD HOSTED
API endpointSelected model

Inference runs on infrastructure we operate.

RESPONSE ↓
Back to your application
Hosted processing · Use with approved workloads
HOSTED BY CUECLOUDOPEN MODELSAPI + CUECODE

CONNECT YOUR WORKFLOW

One endpoint. Your tools.

YOUR CLIENT → HOSTED INFERENCE

Make your first request.

  1. Get access.

    Request hosted access. Once enabled, get your API key from the console.

  2. Configure your client.

    Use the base URL and the model ID provided for your enabled workspace. Keep your key out of source control.

  3. Check the response.

    Try a small, approved coding task before connecting a repository or agent workflow.

Full quickstart ↗
REQUEST / CURL
curl https://api.cuecloud.io/v1/chat/completions \
  -H "Authorization: Bearer cue_••••••••" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_ENABLED_MODEL_ID",
    "messages": [
      { "role": "user", "content": "fix the failing test" }
    ]
  }'

Example only. Replace the key and YOUR_ENABLED_MODEL_ID with your enabled workspace values.

HOSTED MODEL CATALOG

Choose the model for the work.

View catalog ↗

Selected open-weight models offered through our hosted service. Access depends on your enabled workspace; this catalog is not a live capacity monitor.

SELECTED MODEL / HOSTED ON CUECLOUD

GLM 5.3

Hosted inference for your applications and coding workflows.

Deployment
CueCloud hosted inference
Pricing
50% below comparable hosted rates
Access
Enabled workspace required

Your API model ID, request limits, and applicable rates are confirmed during onboarding.

HOSTED INFERENCE / LOWER COST

The models you want.
50% less to run.

How pricing works ↗

Inference for selected open-weight models at 50% less than comparable hosted rates from model labs and neocloud providers.

SAME MODEL / SAME WORKLOAD

What would you save?

$100$10,000

Enter a comparable hosted cost for a selected model. This illustration applies the 50% price difference; it does not quote a live provider rate or set billing terms.

ILLUSTRATIVE CUECLOUD COST
$500/ month in this example

$500 less for the same comparison period.

Match the model version, input/output mix, caching, and service terms. Eligibility and the comparison rate are confirmed with your team.

CHOOSE WHERE INFERENCE RUNS

01 / HOSTED

On CueCloud.

Connect to our endpoint. We run the model infrastructure; your team manages the application and decides which workloads can use the service.

Explore the hosted API ↗
02 / ON PREMISES

Inside your environment.

Vault supplies local inference. Blitz brings the agent workspace, CueCode IDE, execution harness, and telemetry governance into your private environment.

Explore Vault and Blitz ↗

YOUR NEXT STEP

Ready to connect?

Request hosted access and tell us about your workload. We’ll follow up on availability and onboarding. Creating a console account is a separate step and does not itself grant inference access.

Read the quickstart ↗

Create a console account
Already have an account? Sign in ↗