Hardware passthrough Weights on the array Cluster-native scheduling

Serve models on the metal, not around it.

Inference economics come down to two numbers: how much of the machine you can hand to the model, and how fast the weights land in VRAM. Unimatrix0 optimises both. Everything else on this page follows from those two.

99%
Host RAM to the model
A fixed ~2 GB hypervisor. No reserved control-plane pool scaling with the fleet.
100%
VRAM to the model
The physical card, passed through. No virtual framebuffer in the path.
<5s
Cold start from cold array
Weights attach as read-only block and load straight into VRAM.
2s
Model failover
Standby card claimed via disk lease; the endpoint resumes the queue.

Design targets for a reference deployment. We publish measured numbers only after measuring them on your hardware and your model.

Memory economics

The control plane is competing with your context window for RAM.

On a converged stack, every host reserves 16–32 GB for control-plane daemons, container runtimes and shims. On a 256 GB node serving a 70B model at long context, that reservation is the difference between fitting and not fitting.

  • Overhead is fixed, not proportional. It does not grow when you add nodes.
  • Nothing reserves a pool. There is no "spare" memory held back for the machinery.
  • Buy less or serve more. The same host either fits a larger model or one more replica.

How to use this commercially. Model memory is roughly params × bytes-per-param plus KV cache that grows with concurrency and context. Take the RAM you recover, subtract the KV cache you must reserve for your concurrency target, and the remainder is your next partial layer or your next replica. That arithmetic is the whole pitch.

                  256 GB node  ·  70B model  ·  8 replicas, long context

CONVERGED STACK
reserved for control plane    │█████                                             │24 GB   
guest allocator overhead      │█                                                 │6 GB    
model weights                 │███████████████████████████                       │140 GB  
KV cache @ target concurrency │██████████████                                    │72 GB   
                               ✗ does not fit — reduce concurrency or buy RAM

UNIMATRIX0
Dom0 · fixed                  │█                                                 │2 GB    
model weights                 │███████████████████████████                       │140 GB  
KV cache @ target concurrency │██████████████                                    │72 GB   
headroom / 9th replica        │████████                                          │42 GB   
                               ✓ fits, with 42 GB spare

                         Recovered on a 60-node estate: ~1.7 TB.
GPU isolation

The card belongs to a workload, not to the operating system.

Drivers are not installed in the hypervisor. The physical device is assigned to a dedicated guest. That is what makes multi-tenant GPU economics possible — and it is why a bad model cannot take down the machine.

Hard isolation per tenant

Each guest gets a whole card. No time-slicing arbitration, no noisy-neighbour interference, and an accounting boundary that is a hardware boundary.

A crash stays a crash

Driver fault, out-of-memory in the runtime, a pathological kernel: it kills one appliance. The hypervisor, the other tenants and the traditional VMs keep serving.

Density without contention

Eight cards in a host means eight independent failure and capacity domains. Losing one costs you one tenant's throughput, not the node.

Concern Shared-driver model Passthrough to isolated guest
VRAM fidelityVirtual framebuffer overhead100% of the card, native addressing
Driver fault blast radiusHost kernel and every tenantOne appliance instance
Tenant boundarySoftware policyHardware assignment
Runtime upgradeRequires host maintenanceRebuild the appliance; host untouched
Chargeback / meteringApproximatePer-card, per-tenant, exact
Hybrid CPU inferenceMixed paths, mixed overheadSame scheduling model for both
Model delivery

The weights are already on your array. Stop downloading them.

In a container platform, scaling a model server means pulling tens of gigabytes over HTTP per replica — during autoscaling, exactly when your network is busiest. Unimatrix0 keeps weights on shared block storage and attaches them read-only.

40 GB weights · 8 replicas coming online

HTTP pull path
  image layer  ████████████  2–5 min per replica
  + registry round trip
  + autoscaling burst while your WAN is saturated
  × 8 replicas  =  ~320 GB in, nothing cached
  on next deploy: repeat from zero


Block attach path
  read-only attach from shared array
  load into VRAM   ████████████  ~4.8s
  + one-time ingest onto the array
  × 8 replicas  =  same 40 GB on disk, 8 readers
  on next deploy: warm again in seconds


Same bytes. Different fabric.

What this changes operationally

  • Scale-out is local. Adding replica nine does not touch the WAN.
  • Rollback is instant. Previous weights are still on the array; reattach.
  • Air-gapped is the default case. Weights are ingested once, deliberately.
  • Immutable serving. Weights attach read-only, so a served model cannot be edited in place.
  • Versioned by content. Model revisions are array artefacts with names, not image tags.

State it honestly to engineers. This is fast block I/O, not magic. Sustained throughput is bounded by your array's read bandwidth — we size that with you, and we will tell you if your array will become the bottleneck for your replica count.

Scale-out inference

Idle GPUs claim work themselves. There is no scheduler to reconcile.

Inference and batch evaluation requests are published to a shared queue. Each node watches its own VRAM utilisation and thermal state, and an idle GPU claims the work. This is the same mechanism that places virtual machines — one model, not two.

01 / PUBLISH

Submit work, not placement

A batch evaluation or an agent request lands on a cluster-wide queue. You do not choose a node. Nodes are anonymous capacity.

02 / CLAIM

Hardware-aware claiming

Each node evaluates free VRAM, GPU temperature and existing load, then claims only what it can genuinely serve.

03 / EXECUTE

Isolated execution

The job runs inside that card's appliance. Tokens stream back over the bus to the caller. No orchestration layer in the path.

04 / FAILOVER

Survivors take over

If a card's host fails, the workload moves by lease, and the endpoint resumes from the queue rather than dropping the conversation.

Workload shape How Unimatrix0 runs it What makes it efficient
High-concurrency chat / RAG One card per replica, replicas spread across racks Anti-affinity placement, per-tenant VRAM accounting, no shared-driver contention
Batched offline evaluation One queued job, claimed by whichever cards free up first Work-stealing queue replaces an autoscaler and a job controller
Fine-tuning / training Long-running appliances with reserved cards Weights and datasets already on the array; a driver fault is contained
Model hot-swap / canary Attach a different weights artefact, drain the old replica Sub-5 s transition; rollback is a reattach
CPU-only embeddings / small models Same queue, same placement, no card required One scheduling model across CPU and GPU — no second orchestrator
Inference for our own agents On-site sampling capability, fully inside your perimeter Decisions never leave the building; no third-party API in the control path
Agentic operations

The platform can ask a local model for advice, and act on it safely.

Appliances that want help declare that they use the platform's sampling capability. When an AI appliance is present, the request is routed to a model running on your own hardware. When it is absent, those features simply do not appear — nothing breaks.

  • Advice is advisory. It becomes an operation, which then goes through leases, placement and approvals like anything else.
  • Nothing leaves the perimeter. No external inference provider sits in your decision path.
  • Degrades cleanly. No model, no feature. No model, no breakage.
scenario: HA failover completed, cluster looks healthy

platform event
  node pg-07 restarted on umx-7c2e (lease acquired)
  23 sibling restarts in the same batch
        │
        ▼
local model on umx-7c2e  (on your hardware)
        │
        │   reads: logs · leases · node metrics
        ▼
recommendation
  "23 restarts on one surviving host is
   consistent with a rack power event
   and a stale lease cache. Verify r12
   power feed before accepting further
   HA placements into that domain."
        │
        ▼
operator decision
  the model does not place, fence, or approve
  anything. A human accepts the change.
Sovereignty

For a lot of organisations, this is the whole reason.

If your inference has to leave the building, you do not have private AI — you have a contract, a data-processing agreement, and a latency budget you do not control.

No egress

Prompts, completions, weights and telemetry stay on your array. There is no telemetry endpoint, no usage reporting call, no licence server.

Air-gapped capable

Nodes are diskless and discover peers locally. Model weights are ingested deliberately, once. The platform works with no internet route.

Your data lifecycle

Retention, deletion and residency are storage policy, not a vendor's data-handling page.

Your agents

Because the interface is an open protocol, your own models and workflows integrate without a vendor partnership.

Start a conversation

Tell us the model, the concurrency, and the constraint.

Send us the workload profile — model size, target context, peak and sustained request rate, and what the data may not do. We will come back with a sizing model and an honest answer about whether we fit.