State that can go stale
A control database is a second, always-imperfect copy of reality. It drifts from the hypervisor, it blocks on quorum, and it turns a partial network partition into a cluster-wide outage rather than a local one.
Unimatrix0 is a stateless, masterless private cloud for bare metal. It runs your traditional virtual machines, hosts your GPU inference at near-100% memory efficiency, and exposes the entire cluster to AI agents through the Model Context Protocol — on one fabric, with no database to lose.
Boot a node, and it is production. No installer, no controller to deploy, no control plane to lose.
Every mainstream platform puts a stateful database — SQL, etcd, corosync — between you and your hardware. It is the thing that must be patched, backed up, quorum-tuned, and upgraded in a specific order. It is the reason upgrades are events, and why the smallest VM cannot be scheduled if the database is sad.
A control database is a second, always-imperfect copy of reality. It drifts from the hypervisor, it blocks on quorum, and it turns a partial network partition into a cluster-wide outage rather than a local one.
Converged stacks reserve 16–32 GB per host for control-plane daemons, container runtimes and shims. On a 64-core AI node that is memory you bought, paid for, and cannot hand to a model.
When your orchestrator can be driven by an API, but not understood by an agent, every capacity decision still routes through a human reading a dashboard at 3 a.m.
Unimatrix0 does not improve the traditional control plane. It removes it and derives the same outcomes from the hypervisor itself.
Live workload state is read from the hypervisor toolstack. Configuration lives as plain files on shared storage, next to the disks it describes. There is no cache to invalidate, no migration to run, and no backup of the control plane — because there is nothing to back up that is not already on your array.
Nodes are peers, not servers and clients. Any node accepts an API call, publishes work, or claims it. There is no leader to elect, no management role, and no node whose loss degrades the cluster.
Nodes boot a squashed filesystem that never touches local disk and cannot be modified at runtime. A node is disposable: power-cycle it and it is byte-identical. A compromised node cannot persist, and a fleet-wide upgrade is a reboot from new media.
Networking, storage, backup, AI runtimes and protocol parsers all run in isolated guests with their own PCI devices. A kernel panic in a CUDA driver takes down one model server, not the host.
Provision a VM, drain a node, read a lease table, snapshot a workload — each is exposed through the open Model Context Protocol with the same permissions, approvals and audit trail as a human clicking in the console. Every MCP client already speaks it. No SDK, no plugin, no integration project.
Figures below are design targets for a reference deployment on typical enterprise hardware. They describe the architectural ceiling, not a benchmark of your workload.
| Dimension | Kubernetes + KubeVirt | Commercial hypervisor | Unimatrix0 |
|---|---|---|---|
| Control-plane state | Stateful etcd cluster with Raft quorum | Central SQL database, vCenter | None. Hypervisor + storage are the only state |
| RAM reserved per host | 16–32 GB | 4–8 GB | ~2 GB fixed Dom0 — the rest is workload capacity |
| Usable host memory | Reduced; reserved by daemons | Reduced; reserved by the stack | >99% to workloads and model context |
| VRAM available to models | Partial — shared with host drivers | Partial — vGPU layer overhead | 100% via hardware passthrough |
| Node cold boot | 3–5 min | 2–5 min | 18–25 s from live RAM |
| Failure detection | Seconds; depends on health probes | Heartbeat + storage heartbeats | ~2 s signal, then disk-level fencing |
| Split-brain protection | Raft quorum | Storage heartbeats | Atomic disk leases — two hosts can never own one disk |
| 40 GB model load | 2–5 min over HTTP | Moderate | <5 s — SAN block mount into VRAM |
| Agent / AI interface | CRDs + bespoke controllers | Proprietary API | Standard MCP — tools, resources, prompts |
| GPU driver isolation | Host kernel | Host / vGPU layer | Isolated guest with the physical card |
| Upgrade mechanism | Rolling component upgrades, ordering-sensitive | Rolling vCenter + host patches | Drain, reboot from new image, rejoin |
| Persistence after compromise | Attacker can persist on the host | Attacker can persist on the host | None. Reboot eradicates runtime changes |
Reference deployment: 3 nodes, dual-socket Xeon, 256 GB RAM, 2 GPUs per node, 10 GbE iSCSI SAN. Measured behaviour varies with hardware and workload; we publish our own numbers only after we have measured them on your profile.
Inference economics are decided by two numbers: how much of the machine you can give to the model, and how fast you can get weights into VRAM. Unimatrix0 is engineered around exactly those two numbers.
SHARED BLOCK STORAGE (SAN / NVMe-oF) model weights .safetensors .gguf · read-only attach ──────────────────────────────────────────────────── │ block attach │ block attach ▼ ▼ ┌─ umx-node-01 ─┐ ┌─ umx-node-04 ─┐ │ Dom0 2GB 99% free │ │ Dom0 2GB 99% free │ │ sys-ai H100 #0 71% │ │ sys-ai H100 #0 12% │ │ sys-ai H100 #1 88% │ │ sys-ai A100 #0 31% │ └────────────────────────┘ └────────────────────────┘ │ │ │ tasks.ai.global (queue) │ └──────────────────────────┬───────────────────────────┘ ▼ node-04 claims the job — VRAM is free. No scheduler. No YAML.
Point any MCP client at a node. It discovers tools, reads live state, and runs operations — under the same policy engine that governs the web console and the CLI.
unimatrix_create_vm, unimatrix_drain_node, ai_assign_workload. Annotated read-only or destructive, so clients know what they are about to do.
xen://node/domains, sanlock://leases, gpu://inventory. Live cluster state as addressable URIs an agent can read.
Reusable procedures shipped with the product: root-cause a failed failover, plan a restore, rebalance placement — as structured prompts, not tribal knowledge.
Read scopes run freely. Destructive calls raise an approval a human accepts. Every call is attributed, validated against a schema, and written to an audit log.
Growth is a topology problem, and topology is explicit. Nodes carry site, zone and rack labels. Placement spreads replicas across fault domains, reserves capacity for the worst plausible loss, and refuses to take on work it could not survive.
site: eu-west-1 ├── cluster eu-west-1a — primary │ ├── rack r11 node-01 node-02 ♥ leader │ ├── rack r12 node-03 node-04 │ └── rack r13 node-05 node-06 (drained) │ │ placement: pg-primary → anti-affinity: rack │ reserve: N+1 rack ................. satisfied │ lose rack r12 → restart 48 VMs? yes, 62s │ lose rack r13 → restart 48 VMs? no — not enough headroom │ └── cluster eu-west-1b — recovery connected by WAN gateway · autonomous if the link drops
Tell us the model, the concurrency, the recovery objective, and the constraints — air-gapped, sovereign, regulated. We will tell you honestly whether we are the right fit, and what a pilot would look like on your hardware.
No signup wall. You will speak to the engineers who build the platform.