PRIVATE INFERENCE / HARDWARE + RUNTIME

Vault /

Your AI. Plugged in.

On-prem Mac Studios and CueCloud’s custom inference stack, configured for your models and team. Aether-1, our proposed Memory Processing Unit, is the next hardware generation in development.

Plan your Vault deployment ↗
VAULT / SYSTEM ANATOMY ON PREMISES
01 / PHYSICAL SYSTEMCONCEPT VIEW
CUSTOMER ENVIRONMENT04030201COMPUTEExecutes the modelMEMORYWeights + contextLOCAL NETWORKConnects your toolsINSTALLATIONConfigured by CueCloud
01 / Compute executes the model
CUECLOUD / VAULT

Compact compute. Purpose-built deployment.

Consumer-grade hardware, sized for your models and installed by our team.

System concept / configuration varies by deployment

VAULT / THE HARDWARE EVOLUTION

Vault hardware.
Two generations.

Gen 1 uses Mac Studio with our inference software. For Gen 2, we’re developing Aether-1, our own Memory Processing Unit.

VAULT / GENERATION EXPLORERLOCAL INFERENCE SYSTEMS
GEN 1 / SYSTEM ANATOMYAPPLE SILICON
Mac Studio shared-memory inference pathCueCloud inference software schedules work on an Apple Silicon CPU and GPU accessing unified memory. Applications connect through a local inference endpoint. Logical architecture, not a chip floorplan.01 / SHARED MEMORY + CUSTOM INFERENCECUECLOUD INFERENCE STACKModel execution · memory management · schedulingAPPLE SILICON / LOGICAL VIEWCPUHost + orchestrationGPUModel computationUNIFIED MEMORYModel weights + writable state in a shared poolMAC STUDIO / INSTALLED ON YOUR PREMISESLOGICAL PATHS / NOT A PACKAGE FLOORPLAN / NOT TO SCALE
SCHEMATIC / SWIPE TO EXPLORE ↔

The same memory serves CPU and GPU. Serving capacity depends on the model, context and concurrent workload.

CURRENT DEPLOYMENT PATH

Software optimized for Apple Silicon.

Apple Silicon Mac Studios, installed on-site with CueCloud’s custom inference stack. We configure the hardware, models and local serving around your team’s workloads.

WHY THIS GENERATION

Memory shared by CPU and GPU

The CPU and GPU access a common memory pool. Our inference software manages model execution, memory and scheduling within the hardware’s capacity and bandwidth.

  • Apple Silicon CPU + GPU
  • Shared unified memory
  • CueCloud inference + on-site installation
THE DESIGN STEP

Gen 1 runs on Apple Silicon. For Gen 2, we’re developing both the processor and the software that runs on it.

GEN 2 / WHY WE’RE BUILDING AETHER

Why build our own chip?

Running AI on your premises means working with the power, cooling and space you have. We’re building Aether to run larger models within those limits.

When a model exceeds an accelerator’s memory, adding devices can become necessary just to hold it. You pay for the extra compute, then supply the electricity and cooling it needs. We think inference hardware should let you add model capacity without tying every increase to more compute.

OUR DESIGN PRINCIPLE

More room for models should not require the same increase in compute hardware.

01 / CAPACITY

Hold larger models.

Model weights must be available when a request needs them. Keeping more weights in memory does not always require proportionally more computation. We want to use that distinction to support larger models and keep several models available on-site.

02 / BANDWIDTH

Read data fast enough.

The processor needs a steady supply of model weights and request data. If memory cannot supply them fast enough, answers slow down. Capacity and bandwidth have to grow together to make larger models practical to use.

03 / POWER

Stay within the site’s limits.

An office or warehouse has a power and cooling budget. That budget must cover the computers, memory and networking needed to serve your team. We’re designing Aether around those costs and will measure them across the whole system.

What we can change with our own processor.

Mac Studio gives our inference software a memory pool shared by the CPU and GPU. That is the basis of Gen 1. Software can improve how we use the machine, but it cannot change how much memory it has or how quickly the hardware can read it.

With Aether-1, we can choose the balance of memory capacity, bandwidth and computation ourselves. We’re developing the processor and its software together for local inference, with larger models and lower deployment power as our goals. The implementation is proprietary.

Aether-1 is in development. We still need to test capacity, response speed and power together under realistic workloads. Results will depend on the model, context length and number of simultaneous requests. We have not announced product specifications or a release date.

Vault / INCLUDED IN THE DEPLOYMENT

Open the box.
We’ve done the groundwork.

Hardware, software, installation, and handoff.

COMPUTE / MEMORY / LOCAL NETWORK

The hardware

Compact inference hardware, sized for your models and team. We plan the power, networking, and installation with you.

  • Sized around your work

    Compute and memory matched to the models you want to run, the context they need, and the number of people using them.

  • Planned for your site

    We work through available power, cooling, physical placement, and local networking before installation.

  • Delivered as a system

    The hardware arrives as part of a configured inference deployment, with the supported configuration agreed up front.

WHAT YOU WALK AWAY WITH

A configured local inference system sized for your team and ready for installation.

MODEL SERVING / BUILT INTO VAULT

The models you need.
The stack to run them.

Vault is built for inference: running trained models to answer requests from your applications. We supply the hardware and our complete inference stack, licensed to your organization and configured as one system.

YOUR APPLICATIONS / BLITZ / INTERNAL TOOLS

REQUEST ↓   ↑ RESPONSE
01 / SELECTED OPEN-WEIGHT MODELS

Configured for your work

Model selection · Installation · Validation

QwenDeepSeekGLM
Model families to explore during sizing. Specific checkpoints depend on hardware, runtime compatibility, and model license.
02 / CUECLOUD INFERENCE STACK

Included. Installed. Licensed to you.

Model execution · Memory management · Scheduling · Local serving

Our runtime is part of your Vault deployment. You receive a license to use it under your agreement; model weights retain their own license terms.
03 / YOUR HARDWARE CONFIGURATION

Compute and memory, matched.

Compact footprint → More model capacity → More concurrent work

START WITH WHAT YOU WANT TO RUN

Models for the engineering loop.

Code navigation, proposed changes, tool calls, and test feedback put different demands on a model. We help select and configure models for the way your engineers and agents actually work.

  • Code quality on your repositories
  • Tool-use behavior and response format
  • Latency across repeated agent steps
WE HANDLE THE MODEL SETUP

We install the agreed model artifacts, configure supported formats and quantization, set context and concurrency limits, and connect your applications to the local endpoint. Then we validate the configuration on representative tasks with your team.

Explore models and comparisons ↗
A SMALLER FOOTPRINT

Start with the work at hand.

Choose hardware around a focused workload and the models that fit its memory budget. We tune the serving configuration for your team’s expected use.

LARGER MODELS & CONTEXT

Make room for more.

Model weights, working context, and runtime overhead all need memory. We size them together and assess supported hardware arrangements before committing to a configuration.

MORE CONCURRENT WORK

Plan capacity for the team.

More simultaneous requests can call for a different compute layout or separate serving capacity. We customize the deployment around demand, latency goals, power, and placement.

Bring us the models you want to run, or the tasks you need them to do. We help choose the configuration and set it up. Training and fine-tuning are separate requirements to scope.

01 / THE PHYSICAL FOOTPRINT

Power, cooling
and placement.

We check power, airflow, placement and service access before installation.

VAULT / SITE ENGINEERING
01 / POWERCONCEPT SCHEMATIC
VAULT / 01COMPACT ENCLOSURESITE-SPECIFIC LAYOUTLOCAL NETWORKSITE POWER → SYSTEM LOAD

Illustrative enclosure and airflow. Final equipment layout and clearances follow the selected hardware.

Plan the load before the hardware arrives.

We review the selected hardware, available power, operating schedule, and any UPS requirements with your facilities team. We account for both average consumption and peak demand.

  • 01Hardware power requirements
  • 02Available site power and UPS scope
  • 03Measured draw under your workload

EXPLORE THE FOOTPRINT

What changes
as you scale?

Adjust an example load. These are planning inputs, not Vault specifications or measured performance.

AVERAGE COMPUTE LOAD0.30 kW
30-DAY COMPUTE ENERGY72 kWh
APPROX. HEAT INTO THE ROOM0.30 kW thermal

Energy = units × average watts × operating hours × 30 ÷ 1,000. Electrical input becomes approximately the same heat load in the room. Excludes idle time outside the selected hours, networking, UPS losses, and room cooling. Circuit and cooling sizing also require peak-load and site information.

02 / YOUR OPERATING BOUNDARY

Place the intelligence
inside your zone.

Choose the deployment pattern closest to your environment. We work with your IT and security owners to define the actual connections, dependencies, and access paths.

YOUR ORGANIZATIONINTERNAL SERVICE
ENGINEERING TOOLSBlitz / CueCode
INTERNAL APPLICATIONSYour approved clients
REQUEST ↓   ↑ RESPONSE
LOCAL INFERENCE SERVICE
VaultEndpoint → Runtime → Models
Model storageLocal operations
UPDATE + SUPPORT PATHS SCOPED SEPARATELYAdministrative access is agreed with your system owners.

INTERNAL SERVICE

A local endpoint for approved applications.

Your engineering tools and internal applications reach Vault through the network paths your team approves. We scope endpoint access, identity integration, and operational dependencies with your IT owners.

  • Endpoint naming and routing
  • Client access and authentication
  • Agreed update and support paths

We document the approved paths and validate them with your system owners during setup.

03 / SETUP IS PART OF THE SYSTEM

We bring it up.
We show you how it runs.

You should not have to assemble a model-serving project from parts. Our team handles the agreed setup, with your facilities, IT, and workload owners involved where their environment requires it.

01

Scope the work

You bring representative tasks, preferred models, expected users, and the zone where the system will run. We turn those into a hardware and deployment plan.

Deliverable / agreed configuration + site checklist
02

Prepare the system

We assemble the selected configuration, install the inference stack, and configure the best-fit open-weight models for its compute and memory budget.

Deliverable / hardware + runtime + model configuration
03

Install in your zone

We coordinate with your IT and facilities owners, connect the system to the agreed power and network, and configure the endpoint and application connections.

Deliverable / installed system + connection details
04

Validate and hand over

We exercise your representative workload, check operating behavior, and walk your team through everyday use, restart procedures, and the agreed support process.

Deliverable / validation results + operating runbook

AFTER INSTALLATION

A system your team can operate.

We agree how model updates, maintenance, incident triage, and support access will work. In a restricted or offline zone, those procedures are part of the deployment design.

  • Configuration and endpoint details
  • Validation results and operating limits
  • Startup, shutdown, and recovery procedures
  • Named owners and agreed support process

THE COMPLETE SYSTEM

How Vault and Blitz
work together.

Vault runs the models. Blitz puts them to work.
Your organization controls the environment.

YOUR ENVIRONMENT CODE / MODELS / TRACES / TELEMETRY
01 / BLITZAgentic engineering

Your engineers work in Blitz. CueCode IDE is included.

CueCode IDE Write and review code ↗
Agent workspace Direct parallel agent work
Execution harness Run tools and validate changes
Telemetry governance Trace and oversee agent activity
Internal integrations Your repositories, tools, and policies
Explore Blitz ↗
LOCAL
INFERENCE
02 / VAULTLocal inference

Run the models behind Blitz on hardware inside your environment.

  • Compact hardware
  • Custom inference stack
  • Open-source models configured for your system
Explore Vault ↗
↳ HOSTED AND OPERATED INTERNALLYAIR-GAPPED CONFIGURATION AVAILABLE TO SCOPE

For an air-gapped installation, model access, identity, updates, and support are configured and validated within the agreed boundary.

Plan your
Vault deployment.

Tell us about your team and workload. We’ll define the hardware, software, and support scope with you.

Start a conversation ↗