The Lab
Six systems. No switch.
Everything measured.
JMNI Labs runs a working private-inference lab built on NVIDIA GB10-class hardware. It exists for one reason: nothing ships from here without numbers behind it.
The fleet
- Six GB10 systems — NVIDIA DGX Spark and ASUS GX10 machines, 128 GB of unified memory each.
- Switchless rings. Nodes are cabled directly over their ConnectX-7 links in 2- and 4-node rings — collective traffic rides the NICs' own hardware forwarding, so every node can reach every other without putting a switch in the model-traffic path.
- Multi-node tensor-parallel serving for frontier open-weight models, quantized to NVFP4 where it makes sense.
- Everything documented. Each configuration ships with a build log, a benchmark record, and a reproducible deployment kit.
What runs here
GLM-5.3-Flash
NVFP4 · TP2
Tensor-parallel serving on a DGX Spark pair, with the deployment kit published alongside its measured results.
Qwen3.8-Flash-Next
NVFP4 · TP1 / TP2
Single-Spark and pair deployments — including a hybrid checkpoint whose mixed precision beats the stock release on GB10.
DeepSeek-V4.1-Flash
TP4 · SWITCHLESS RING
Served across a four-node ring (2× DGX Spark + 2× ASUS GX10) with full benchmark documentation.
How we work
- Build. We stand the system up on our own hardware first.
- Measure. Every deployment gets a benchmark record — decode rates, prefill, concurrency — not vibes.
- Publish. Kits, tools, checkpoints, and field notes go out in public. If it doesn't hold up in the open, it doesn't ship.
Working with GB10 systems, or planning to? We're taking design partners for our deployment kits and subscription bundles. Get in touch →