Two GB10 systems become one API on port 8888
This kit serves DeepSeek v4.1 Flash through vLLM's OpenAI-compatible interface, split with tensor parallelism across two NVIDIA GB10 DGX Sparks. Its EXL3 checkpoint averages 2.9 bits per weight, while native Engram tables stay outside that quantized tree. DSpark speculative decoding uses draft experts already inside the checkpoint. There is no separate draft model to fetch or operate.
The hardware target is fixed rather than suggested. The container is Linux ARM64, its native code targets sm_121a, and the two ranks communicate over CX7. The server exposes model id DeepSeek-v4.1-Flash-EXL3 on port 8888. If your machines, network interfaces, or CUDA architecture differ, this repository is the wrong starting point unless you intend to maintain the port yourself.
About 387 GiB of model data sits outside the small checkout
The download is split between roughly 197 GiB of EXL3 shards and about 190 GiB from two original DeepSeek shards that hold Engram data. A clean recipe can pull a published image of about 9 GiB instead of compiling the overlay. The head then shares or copies weights to the worker through NFS, rsync, or an optional ZFS path.
That layout explains why setup is operational work, not docker run. You need working SSH, user and group IDs, Docker on both nodes, the correct CX7 interfaces, and enough space for the chosen sharing mode. The default NFS route avoids a second full model copy. ZFS needs about 385 GiB free on both nodes, which the README says the example worker does not have.
What happened when we ran it
Our sandbox installed commit 6f7d159 in 21 seconds. It added 35 Python packages and occupied 37 MB. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. The repository contained 96 files, roughly 14,110 lines of source, a Dockerfile, and a tests directory. It had no CI workflow files.
The build succeeded in 8 seconds. Pytest failed after 9 seconds, before running a test: 0 passed, 0 failed, and all 3 reported items ended as collection or setup errors. The log tail shows tests/test_row_store.py raising SystemExit because librow_store.so was missing, followed by the instruction to build the serving image first. That is the stated failure, and we did not run the two-Spark service.
Pip-audit reported 0 known vulnerabilities in the 35-package environment we installed. That result covers those Python dependencies, not the published container, CUDA code, model files, Docker host, or optional native extension. Likewise, the successful 8-second build does not prove that weights load, ranks connect, or an inference request completes on GB10 hardware.
A few GiB of headroom carries the serving risk
The README budgets about 99.5 GiB of EXL3 weights per rank, a 2.5 GiB KV pool, several GiB for contexts and CUDA workspaces, and more for the operating system and Docker. Its memory guard is disabled by default. The documented reason is uncomfortable: the guard once killed the serving container when an unrelated host upload consumed memory, even though vLLM was idle.
Other controls include a 12 GiB boot-margin check, allocator release after prefills, page-cache dropping, and a high OOM score for the containers. These are thoughtful responses to unified CPU and GPU memory on GB10. They are also evidence that this service should have dedicated nodes. A casual upload or desktop workload should not compete with hundreds of GiB of model state.
Open reports reach restart, first boot, and image requests
The repository was created on September 13, 2026, last pushed on September 19, and had 249 stars on October 3. GitHub counted 16 open issues and pull requests. Issue activity continued through September 26, so the lack of a later code push is not enough to call the project abandoned. There is no GitHub Release, despite the published container image.
Issue 1 reports a missing Engram configuration file and an NFS export failure on a fresh two-node install. Issue 26 says the rsync preflight ignores an existing worker copy and can demand another 385 GiB on restart. Issue 28 reports one large multi-image request stalling at the shipped 1,536-token chunk. Issue 17 says the optional cooperative MoE guide names a pinned binary that operators cannot obtain.
Those reports do not erase the engineering in the recipe. They define its buyer. This is for someone who owns the exact two-node kit, reads every preflight check, and can debug NFS, CUDA memory, and request scheduling. Our 21-second install is the smallest part of that job; the real product is an appliance you will operate yourself.

