Products
Built on a working lab,
not a slide.
Everything we ship is proven on our own six-system GB10 fleet first — then released as a kit, a tool, a checkpoint, or an application.
Local AI infrastructure
Sparkdeck
DGX SPARK FLEET CONTROL
Unified control for a fleet of DGX Sparks: launch, manage, and stop inference workloads across your systems from one place.
sparktop
RESOURCE MONITOR
An rtop-style monitor for DGX Sparks — live GPU, memory, and host telemetry for your lab, in the terminal where you live.
Deployment kits
REPRODUCIBLE MULTI-NODE SERVING
Benchmarked, documented deployment packages that stand open-weight models up on GB10 hardware — single-node and multi-node:
- Qwen3.8-Flash-Next — TP1 and TP2 deployments on DGX Sparks, including a published hybrid checkpoint.
- GLM-5.3-Flash NVFP4 — tensor-parallel serving on a Spark pair.
- DeepSeek-V4.1-Flash — TP4 across a four-node switchless ring (2× DGX Spark + 2× ASUS GX10).
Published checkpoints
JMNI-LABS ON HUGGING FACE
When a hybrid mixture of trained tensors and quantized components outperforms stock releases on GB10-class systems, we publish it ourselves:
JMNI-Labs/Qwen3.8-Flash-Next-NVFP4-QAD5500-Hybrid — a 93B image-text-to-text hybrid for Spark pairs.
Subscription kits — in development
FOR CUSTOMER-OWNED CLUSTERS
Model packs, deployment kits, and rolling updates for clusters you own — designed so a small team can operate a private inference stack without becoming a full-time ops crew. Currently in testing with design partners.
Recitation
Your personal university professor.
Upload any textbook — PDF, txt, or markdown — and Recitation builds a complete course around it: an auto-generated curriculum, interactive streamed lectures, graded problem sets, and a professor that remembers what you've learned and what you already know.
- Grounded in your book. Answers cite the pages they came from. Recitation searches the indexed book first, supplements with general knowledge, and never invents page numbers.
- Voice in and out. Speak your questions (local whisper transcription); hear lectures in human-sounding voices, generated on your machine (Kokoro / Supertonic).
- Runs entirely on your hardware. Point it at your own cluster, Ollama, LM Studio, or any OpenAI-compatible endpoint. Fully offline-capable once models are on disk.
- Any model works. No native function-calling required — Recitation supports weak local models as gracefully as frontier ones.
Want it early, or need it for a school or organization? Talk to us.