The questions we get every time.
Including the awkward ones. If something is not here, ask — and if the honest answer is "not yet", we will say that on the page too.
Platform
What exactly is Unimatrix0?
A private cloud platform for bare-metal servers. Each node is a small, immutable operating system with a mature hypervisor, connected to shared storage and to its peers. There is no central management server and no database. Any node can accept an administrative request, and the cluster decides where work runs.
It is built for organisations that want infrastructure they control — sovereignty, predictable economics, and no dependency on a vendor's roadmap.
How is this different from just running KVM on a Linux distribution?
A single server running KVM is not a cluster. The moment you add hosts you need answer to questions about failure detection, which host may start which workload, and how two hosts are prevented from writing the same disk.
Unimatrix0 answers those three questions by construction: failure is detected on a short heartbeat, placement is derived from live hypervisor state, and mutual exclusion is an atomic claim on the storage array rather than a quorum that can be lost. You do not assemble a control plane to get those properties — they are the architecture.
Why no database? Most clusters seem to need one.
They need it because their architecture treats the management database as the authoritative record of what exists. In Unimatrix0 the authoritative records are the hypervisor and the storage array, which are already authoritative.
Live workload state is read from the hypervisor directly. Configuration lives as plain files on shared storage. There is no second copy of reality to keep in sync, which means there is no drift, no failover of the database, no quorum to lose, and nothing extra to back up.
What happens if the management node goes away?
Nothing — there isn't one. Every node is identical and can serve administrative and agent requests. Losing any single node, including the one you were logged into, has no effect on the cluster's ability to be managed. Losing several leaves reduced capacity and, eventually, workload failure — but never a control-plane outage.
What hypervisor is this built on?
Xen, in a deliberately minimal domain zero. Xen is type-1, has years of production use in enterprise hosting, and has mature hardware passthrough for GPUs — which is what makes the AI story work. It also has a small, well-understood control surface.
Guest operating systems are unmodified. Windows Server, mainstream Linux, BSD and purpose-built appliances run without special guest agents.
What guest operating systems are supported?
Standard Linux distributions, BSD, and Windows guests through device-model emulation. Network and storage peripherals are presented through industry-standard virtual hardware that mainstream guests have drivers for.
Guests with legacy or vendor-specific hardware requirements can be attached with PCI passthrough, so a guest that needs a specific device can have the actual device.
Can I keep my existing VMs and storage?
Yes. Existing disks on shared storage can be attached directly, and workloads are rebuilt against the cluster rather than re-imaged. The migration is a rebuild plus verification pass per workload — typically hours per node, not days.
A cold-migration path for moving a workload between separate clusters is also available, for consolidating estates or moving between sites.
Which storage backends are supported?
Shared storage is the foundation, and block storage is what we recommend for production: Fibre Channel, iSCSI and NVMe over Fabrics are all supported, and any storage exposing raw LUNs works. NFS is supported and useful for smaller deployments and for the image library.
Multiple backends can be registered and used simultaneously, so you can tier by performance or residency — fast local arrays next to capacity storage, or separate backends per tenant.
AI & Inference
How do you get close to 100% of host memory for workloads?
By not running a control plane on the host. Each node's hypervisor domain has a small, fixed allocation — the same on a 32-core node and a 128-core node. It does not grow with fleet size, and nothing reserves a pool against it.
Everything above that line is available to workloads and model context. On a 60-node estate that is on the order of 1.7 TB of memory that a conventional stack would have reserved for its own machinery.
Which models and runtimes are supported?
Anything that runs on a normal Linux server with access to the accelerator, because an AI appliance is a normal guest with the physical card handed to it. That covers the major serving runtimes and training frameworks out of the box, plus anything containerised.
We are opinionated about the packaging and scheduling around it — weights on shared storage, claimed work, health checks — and deliberately agnostic about which framework is inside.
Can I mix CPU-only and GPU workloads?
Yes, and they share one scheduling model. Embeddings, small models, classification and classic virtual machines all sit in the same queue system as large GPU inference. There is no second orchestrator to learn and no capacity planning across two systems.
How does scale-out inference work without an autoscaler?
Requests are published to a shared cluster-wide queue. Each node watches its own accelerator utilisation and thermal state, and an idle node claims the work it can genuinely serve.
This is work-stealing rather than a scaling controller. There is no replica-count field to tune, no scale-to-zero heuristic to misfire, and no second control loop that has to agree with the first. You publish work; capacity finds it.
Is multi-tenancy safe on a shared GPU host?
Each accelerator is assigned to a specific guest at the hardware level. Tenants do not share a card, time-slice a driver, or contend for a scheduler inside a hypervisor. The isolation boundary is the device itself.
If a tenant's workload faults the driver, it faults inside their guest. The other tenants on that host, the hypervisor, and the traditional workloads are unaffected — and the per-card accounting makes it straightforward to charge back.
How do you serve models at high concurrency without memory pressure?
Replicas are placed across distinct accelerators with anti-affinity, and you choose the replica count. Because there is no reserved memory pool on the host, more of the machine is genuinely available for the weights and the key-value cache.
We will size this with you honestly — including telling you when your concurrency target does not fit in the memory you have, rather than discovering it in production at peak.
What happens to inference when a GPU node fails?
The failure is detected on the short heartbeat, the stale storage claim is released, and the workload is placed on a surviving node with a standby accelerator. The endpoint resumes from the queue rather than dropping the conversation.
Recovery time depends on your hardware and storage latency. We will measure it on your equipment and give you the real number.
Can the platform run inference for its own agents on-premises?
Yes. Appliances can declare that they want help deciding something, and the request is routed to a model running on your own accelerators. Nothing leaves the building.
The output is advisory. Anything it suggests still goes through the normal operation path — storage claims, placement policy and human approval. A model never bypasses the safety plane. And if no model is deployed, those features simply are not offered.
Agents & MCP
What is MCP and why does it matter here?
The Model Context Protocol is an open standard for how an AI application discovers and uses external capabilities: things it can do, things it can read, and reusable procedures.
It matters because it means agents operate infrastructure through a declared, discoverable interface with typed inputs — rather than by guessing at shell commands. Every major assistant and agent framework already supports it.
Which agent clients are supported?
Any client that speaks the protocol. It is reached over the standard HTTP transport on any node, so there is no special endpoint and no vendor-specific client to obtain. If it implements the protocol, it works.
Can I restrict what an agent is allowed to do?
Yes, at three levels. Roles map to permission patterns. Tokens are scoped to a role. And a session only sees the tools its role permits — a tool it may not call is not advertised to it at all.
Destructive operations additionally require a human approval, and where the client supports it, the request surfaces in the user's own interface for a decision.
What stops an agent from doing something damaging?
Three independent mechanisms:
- Least privilege. The agent only sees what its role allows.
- Schema validation. Inputs are validated against declared schemas before anything executes.
- Approval gates. Destructive operations stop and wait for a human.
Underneath all of it, storage safety is enforced independently of any caller. Even a fully-permissioned agent cannot cause two hosts to write one disk, because that is decided by the storage array, not by the request path.
Can I review what an agent did?
Every operation records the caller, the surface it arrived on, the tool, an argument digest and the outcome. That record is available in the console and exportable for your audit process.
This is the capability we are asked for most often in agentic evaluations, and the one most systems built around agents do not have.
Do I need a gateway to expose it to agents outside my network?
No, and we would generally advise against it. The endpoint is management-network scoped by default. If you need access from a broader network, an optional hardened gateway appliance handles TLS termination, rate limiting and tenant policy.
In most deployments the right answer is that the agent runs inside your perimeter, which is also the answer to most security questionnaires.
Can I wrap an MCP server we already use?
Yes. A third-party protocol server can be wrapped as an ordinary extension. Its capabilities are declared in a manifest, which means they get consistent names, the same permission model, the same audit trail and the same approval behaviour as native operations.
Common candidates are network controllers, storage vendor tooling and internal service catalogues.
Does adding an extension require writing an integration?
No. An extension declares its capabilities once in a manifest. The platform generates the API endpoint, the command, the console screen and the agent tool from that single declaration.
So every capability you add is immediately available to agents, with no separate integration work — ever.
Security
What is the attack surface of the hypervisor itself?
There is no exposed privileged service on the network. The hypervisor domain has no public interface and no unauthenticated socket. Every request must enter through an application that is designed to be the entry point.
An attacker must first compromise a workload or a boundary appliance to reach anything of value — and that compromise is contained in a guest, not the control plane.
If an attacker gets onto a node, what stops them persisting?
The operating system is a read-only image running from memory. Runtime changes to the system do not survive a reboot, because the next boot restores the exact signed image.
There is no "re-image the host" procedure, because there is no install. Power-cycling the machine is the entire remediation.
Do you have a security certification?
Ask us directly and we will tell you precisely: which frameworks we have been assessed against, which we are working toward, and which we do not hold.
Where you have an obligation requiring attestation we cannot yet evidence, we would rather surface that in the first conversation than the fourth. We would also rather help you build the case from architectural properties — immutability, no exposed privileged service, contained appliances, full attribution — than ask you to take our word for a badge.
Does any telemetry leave my environment?
None. There is no usage reporting, no licence activation call, no crash upload and no external dependency in the operation path. Your environment's behaviour is not our business.
This is what makes air-gapped and sovereign deployments straightforward rather than an engineering project.
Can I run this with no internet connectivity?
Yes. Nodes discover peers locally, hold all state on your storage, and ingest model weights and images deliberately through a controlled process. There is nothing that requires an external route to function.
How is split-brain actually prevented?
By an atomic claim on shared storage. Before a node may start a workload, it must acquire an exclusive lease on that workload's disk. Two hosts cannot both hold it — the storage array is the arbiter.
A hardware watchdog is bound to lease renewal, so a host that has lost connectivity is reset before it can appear dead-but-alive to its peers.
Honest limitation: this trades a small amount of recovery latency for correctness. A network problem can delay recovery. It cannot produce two writers, and no amount of quorum tuning would have changed that.
What happens if my storage array is unreliable?
Workloads that are already running keep running, because they hold their data locally in memory and on their own disk. Configuration changes and failovers will stall while the array is unreachable.
No software above the storage layer can compensate for an array that loses acknowledged writes. We size storage with you for this reason, and we will be direct if your array cannot support the safety guarantees you need.
Can I integrate this with our own identity provider?
Role-based access with scoped tokens is the model. Federation with your existing identity provider is on the roadmap and is a common requirement — tell us what you run so it is scoped correctly rather than assumed.
Operations
How much work is it to run this day to day?
The premise of the architecture is that operational load scales with the number of nodes, not the number of workloads. Nodes are disposable and self-healing. Failover is automatic. There is no database to repair and no controller to re-elect.
Our cost model assumes roughly 22 nodes and 110 AI workloads per platform engineer, which is the number to argue with. If your ratio is materially worse, we want to know why before you buy.
How do upgrades work?
Nodes hold no local state, so upgrading means booting a newer image. The routine is: drain the node, reboot it from new media, watch it rejoin, repeat.
Mixed versions during a roll are expected and visible — every node reports the exact release it is running, and the console flags anything out of date. Rollback is booting the previous image.
Do I need local storage on the nodes?
Not for the platform. Nodes can boot with no local disk at all, holding everything in memory. Local storage is optional and, where present, useful for workload data that needs local latency or local-only residency.
Can nodes boot over the network?
Yes, and it is a natural fit precisely because nodes are stateless. A single release pointer on the provisioning server becomes a fleet-wide upgrade, and rollback is flipping it back.
The trade-offs are honest ones: a boot-time dependency on the provisioning server, credential handling on the provisioning path, and bandwidth for each boot. Sites typically mirror locally.
How do I add a node?
Enable hardware virtualisation in firmware, boot the image, and point it at the shared storage. It discovers its peers, joins, and starts accepting production workloads. There is no installer, no configuration management run, and no registration database.
How do I take hardware out of service?
Mark it as draining. Workloads live-migrate off with rate control so you do not saturate the fabric, and new placements avoid it. When it is empty, power it down.
You can also drain an entire rack or fault domain at once, which is the operation you actually want during maintenance windows.
What happens to my current monitoring and backup tooling?
The platform exposes standard interfaces — a REST API, a CLI, a streaming event feed, and an MCP endpoint — so existing automation is straightforward to repoint.
A backup appliance handles snapshots and off-site shipping natively, and snapshots created externally can be picked up by it.
How do I get support?
Every tier gets access to the engineers who build the platform, not a first-line script. Community is forum-based; Standard adds business-hours support with a one-business-day response; Scale and Enterprise add 24×7 with contractual response times and, at Enterprise, a named engineer who knows your estate.
We also run failure drills with you. Know your recovery time before you need it.
What is the fastest way to evaluate this?
Run a pilot on hardware you already own, on a non-critical workload, and let it run through a real failure. Nothing sells this architecture like watching a workload survive the loss of the machine it is running on.
Tell us what you have available and we will help you scope it. Start the conversation.
Commercial
How are you priced?
Annual subscription per node, with every capability included. There is no per-vCPU meter, no workload metering, no storage charge, and no control-plane edition.
The reasoning: metering usage on a platform that removes the metering problem is self-defeating. You pay for the hardware you run the platform on, and that is it.
Is the AI fabric included or an add-on?
The capability is always included — inference, model delivery and the local sampling capability are part of the product at every tier.
Where most of your estate is accelerator nodes, it is packaged per accelerator node at $940/node/year. That is a packaging choice to fit how you buy hardware, not a functional restriction.
How confident should I be in the cost model?
Treat it as a structured argument, not a quote. Every assumption is published on the pricing page so you can replace any of them with your own figures. If you substitute real numbers and the conclusion changes, tell us — it means we have misunderstood your estate.
What about professional services?
Quoted separately at $1,850 per engineer-day. A first deployment is usually 8–15 days depending on node count and integration complexity.
This exists because a cluster that is deployed badly fails in ways that look like product defects. We would rather you paid for the deployment than for the recovery.
What happens to our data if we leave?
You keep it. Configuration is plain files in documented formats. Disks are standard images on your storage. Nodes are disposable, and there is no controller, no licence file and no activation server that could stop booting.
Leaving costs nothing, which is the intended consequence of having no lock-in.
Do you offer a trial?
A Community licence is free and unlimited in duration for up to three nodes. We would rather you ran a real evaluation on your own hardware than a scripted demo in ours.
Question we did not answer?
Send it. If it is a hard question, we would rather have it now than after you have committed an engineering team.