Claims & validation

Which of our numbers we have actually measured.

Almost none yet. This page says which, so that you do not have to guess — and so that we cannot quietly imply otherwise later.

12

Published claims

Every quantitative or correctness claim made anywhere on this site.

2

Verified by construction

Correctness properties of the mechanism rather than benchmark results.

10

Awaiting benchmarks

Stated and argued, not yet measured on reference hardware.

0

Modelled

Derived from a published assumption you can replace with your own figures.

The register

Entries move between states as benchmarks land. We do not remove rows: the distance between stated and measured over time is itself a useful signal.

Register of published claims, their validation state and their basis
Claim State Basis
Domain zero reserves approximately 2 GB per host. Stated, not measured Memory-resident Linux, hypervisor and one daemon. Fixed by design: it does not grow with fleet size or workload count.
Over 99% of host memory is available to workloads and model context. Stated, not measured Follows from the fixed reservation above. Excludes guest-allocator overhead and any kernel reservation, which are not claimed.
Kubernetes + KubeVirt reserves 16–32 GB per host. Stated, not measured Illustrative of a typical deployment. We have not measured this on a competitor platform and do not hold a benchmark.
Commercial hypervisors reserve 4–8 GB per host. Stated, not measured Illustrative of a typical deployment. See the note below — this figure is the least validated entry on the page.
Node cold boot is 18–25 seconds. Stated, not measured Firmware to a schedulable hypervisor, from RAM, with no installer phase. Excludes physical hardware POST and PXE transfer on a cold network.
Failure detection is approximately 2 seconds. Stated, not measured Heartbeat plus disk-level fencing. This is signal-to-decision time, not time-to-restored-service.
A node boots into production with no installer and no controller. Stated, not measured Immutability and statelessness are architectural properties, not performance claims. They are either true of the design or they are not.
100% of VRAM is available to the model. Stated, not measured Hardware passthrough. There is no vGPU or mediated-device layer in the path. Excludes memory the driver itself reserves.
A 40 GB model can start serving in under 5 seconds. Stated, not measured Weights attached read-only from shared block storage and loaded straight into VRAM, rather than pulled per replica over HTTP.
Model weights are shared across replicas from one set of blocks. Stated, not measured The storage and model-delivery architecture. This is the mechanism we think is most commercially interesting, and it is the one most in need of a benchmark.
Two hosts can never own the same disk. Verified by construction Atomic disk leases with fencing. Correctness is a property of the mechanism, not a benchmark. See the security model for what this does not cover.
There is no control-plane database to lose. Verified by construction The hypervisor and the storage layer hold all state. Message bus and peer discovery are rebuilt rather than recovered.

The weakest entry on this page is the "4–8 GB" commercial hypervisor figure. It is illustrative, it is the one most likely to be wrong in a way that flatters us, and a reader who knows their own estate can contradict it in one line. It stays here, labelled, until we measure it properly — because removing it would hide the problem rather than fix it.

What a measurement would look like

The standard we intend to hold ourselves to.

When we publish a benchmark it will carry the full configuration, not a summary. If we cannot get that right for one number, we will not get it right for fifty.

Every published benchmark states

  • Exact hardware, including CPU, GPU and storage controller models.
  • Firmware versions and the umx build identifier.
  • Model, quantisation and context length.
  • Concurrency and request pattern.
  • What was measured, how, and over how long.
  • What the number does not include.

The metric that matters most

Cost per million tokens. It is the only figure that combines hardware utilisation, throughput, memory efficiency and staffing into something a finance team can act on, and it is the one we intend to be judged on.

We would rather publish one honest number on this than forty architectural claims. The cost model on the pricing page is the current approach, and it is a placeholder for this until real measurements exist.