mrkeyoor.com_
Sat 03 Oct 07:18 UTC
LLM Toolsevaluationupdated 03 Oct 2026

DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks review

This repository is a deployment kit for serving an EXL3-quantized DeepSeek v4.1 Flash model across two NVIDIA DGX Spark GB10 systems. It wraps custom kernels, split model storage, vLLM, and an OpenAI-compatible API into one highly specific two-node recipe.

Verdict

Our host build finished in 8 seconds, but pytest stopped after 9 seconds with 3 collection errors because the serving image's librow_store.so was absent, so a cheap generic container cannot validate the real appliance. Use this kit only for its exact job: serving this checkpoint on two GB10 DGX Sparks. For a normal GPU fleet or a service with conservative release gates, start with an upstream runtime and avoid the custom storage, kernel, and memory chain.

We ran it

Lab card: what happened when we ran DeepSeek-v4.1-Flash-EXL3-2x-DGX-SparksScreenshot of DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks (x.com/MiaAI_lab)
Install✓ · 21s35 packages · 37 MB
Build✓ · 8s
Tests✗ · 9s0 passed · 0 failed · 3 errors of 3 (pytest)
Known vulns0(pip-audit)
Repo96 files~14,110 lines of source · 1.1 MB · 0 CI workflows · Dockerfile · tests dir

Answers from our run

Does DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks build from source?

Dependencies installed in 21 seconds (35 packages), and the build succeeded in 8 seconds. We cloned commit 6f7d159 into a clean Debian container with 3 CPUs and no project-specific setup.

Do DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks's tests pass?

Yes: 0 of 3 passed when we ran the project's own test command (pytest), with 3 collection errors. Some failures need services or credentials a bare container does not have.

Does DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks?

Anyone without two GB10 DGX Sparks: the image, native cubins, tensor parallel layout, and network checks target that pair.

What are the alternatives to DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks?

vLLM, SGLang, TensorRT-LLM. Our host build finished in 8 seconds, but pytest stopped after 9 seconds with 3 collection errors because the serving image's `librow_store.

Setup1/5Two GB10 nodes, 387 GiB of weights, and custom networking
Docs5/5Exact memory, storage, kernel, rollback, and failure guidance
Community3/5249 stars with detailed operator reports and active patches
Maturity2/5No release tag; fresh-boot and request-path defects remain open

Who it’s for

Owners of exactly two DGX Spark GB10 systems connected through the expected CX7 path.
Model-serving engineers comfortable with Docker, NFS or rsync, SSH, CUDA extensions, and memory tuning.
Teams that need a local OpenAI-compatible DeepSeek endpoint and accept a specialized runtime.
Researchers willing to validate quantization quality and multimodal behavior against their own workloads.

Who it’s NOT for

Anyone without two GB10 DGX Sparks: the image, native cubins, tensor parallel layout, and network checks target that pair.
Operators short on storage or transfer time: the README calls for about 387 GiB of weights plus an image of roughly 9 GiB.
Teams that need a hardware-neutral container: the kit depends on custom SM121 code, CX7 communication, split Engram data, and two-node coordination.
Release processes that require a green generic-host suite: our pytest run collected 0 passing tests and stopped on 3 errors because librow_store.so was missing.
Operators depending on the optional cooperative MoE path: issue 17 says the pinned binary named by the guide is unavailable, and the repository has no GitHub Releases.
Organizations whose distribution model is incompatible with the launcher's AGPL-3.0 license.

Setup reality

Our fresh Python 3.12 sandbox installed 35 packages in 21 seconds and used 37 MB. The build succeeded in 8 seconds. Pytest failed after 9 seconds with 0 passed, 0 failed, and 3 collection/setup errors; test_row_store.py exited because librow_store.so was missing and said to build the serving image first.

Actual service setup needs two DGX Spark GB10 nodes, Docker, SSH between head and worker, CX7 networking, about 387 GiB of model data, and either NFS, rsync, or ZFS for weight access. The public image needs no registry login; downloads use the Hugging Face CLI.

The shipped memory settings leave little headroom, and the watchdog is disabled by default. Open reports describe fresh NFS boot failures, an rsync restart check demanding another 385 GiB, and a multi-image request livelock at the default token chunk. Pip-audit found 0 known vulnerabilities in our 35-package environment.

Two GB10 systems become one API on port 8888

This kit serves DeepSeek v4.1 Flash through vLLM's OpenAI-compatible interface, split with tensor parallelism across two NVIDIA GB10 DGX Sparks. Its EXL3 checkpoint averages 2.9 bits per weight, while native Engram tables stay outside that quantized tree. DSpark speculative decoding uses draft experts already inside the checkpoint. There is no separate draft model to fetch or operate.

The hardware target is fixed rather than suggested. The container is Linux ARM64, its native code targets sm_121a, and the two ranks communicate over CX7. The server exposes model id DeepSeek-v4.1-Flash-EXL3 on port 8888. If your machines, network interfaces, or CUDA architecture differ, this repository is the wrong starting point unless you intend to maintain the port yourself.

About 387 GiB of model data sits outside the small checkout

The download is split between roughly 197 GiB of EXL3 shards and about 190 GiB from two original DeepSeek shards that hold Engram data. A clean recipe can pull a published image of about 9 GiB instead of compiling the overlay. The head then shares or copies weights to the worker through NFS, rsync, or an optional ZFS path.

That layout explains why setup is operational work, not docker run. You need working SSH, user and group IDs, Docker on both nodes, the correct CX7 interfaces, and enough space for the chosen sharing mode. The default NFS route avoids a second full model copy. ZFS needs about 385 GiB free on both nodes, which the README says the example worker does not have.

What happened when we ran it

Our sandbox installed commit 6f7d159 in 21 seconds. It added 35 Python packages and occupied 37 MB. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. The repository contained 96 files, roughly 14,110 lines of source, a Dockerfile, and a tests directory. It had no CI workflow files.

The build succeeded in 8 seconds. Pytest failed after 9 seconds, before running a test: 0 passed, 0 failed, and all 3 reported items ended as collection or setup errors. The log tail shows tests/test_row_store.py raising SystemExit because librow_store.so was missing, followed by the instruction to build the serving image first. That is the stated failure, and we did not run the two-Spark service.

Pip-audit reported 0 known vulnerabilities in the 35-package environment we installed. That result covers those Python dependencies, not the published container, CUDA code, model files, Docker host, or optional native extension. Likewise, the successful 8-second build does not prove that weights load, ranks connect, or an inference request completes on GB10 hardware.

A few GiB of headroom carries the serving risk

The README budgets about 99.5 GiB of EXL3 weights per rank, a 2.5 GiB KV pool, several GiB for contexts and CUDA workspaces, and more for the operating system and Docker. Its memory guard is disabled by default. The documented reason is uncomfortable: the guard once killed the serving container when an unrelated host upload consumed memory, even though vLLM was idle.

Other controls include a 12 GiB boot-margin check, allocator release after prefills, page-cache dropping, and a high OOM score for the containers. These are thoughtful responses to unified CPU and GPU memory on GB10. They are also evidence that this service should have dedicated nodes. A casual upload or desktop workload should not compete with hundreds of GiB of model state.

Open reports reach restart, first boot, and image requests

The repository was created on September 13, 2026, last pushed on September 19, and had 249 stars on October 3. GitHub counted 16 open issues and pull requests. Issue activity continued through September 26, so the lack of a later code push is not enough to call the project abandoned. There is no GitHub Release, despite the published container image.

Issue 1 reports a missing Engram configuration file and an NFS export failure on a fresh two-node install. Issue 26 says the rsync preflight ignores an existing worker copy and can demand another 385 GiB on restart. Issue 28 reports one large multi-image request stalling at the shipped 1,536-token chunk. Issue 17 says the optional cooperative MoE guide names a pinned binary that operators cannot obtain.

Those reports do not erase the engineering in the recipe. They define its buyer. This is for someone who owns the exact two-node kit, reads every preflight check, and can debug NFS, CUDA memory, and request scheduling. Our 21-second install is the smallest part of that job; the real product is an appliance you will operate yourself.

Alternatives

ProjectWhat it isPick it when
vLLM gh↗A general high-throughput model server with OpenAI-compatible APIs and broad hardware support.pick this instead when your model and hardware already work upstream without this kit's EXL3 and GB10 patches.
SGLang gh↗A serving runtime for language and multimodal models with its own scheduling and kernel stack.pick this instead when you want a supported SGLang deployment rather than this pinned vLLM overlay.
TensorRT-LLM gh↗NVIDIA's inference stack for optimized language-model deployment on supported GPUs.pick this instead when vendor-supported model paths matter more than running this exact quantized checkpoint.

What people are saying

  1. [velocity-scout] MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks

Sources

  1. DeepSeek v4.1 Flash EXL3 deployment README
  2. Repository facts
  3. Fresh two-node boot report
  4. Rsync restart free-space report
  5. Multi-image request livelock report
  6. Missing cooperative MoE artifact report

More llm tools reviews

whatsapp-mcp · deepseek-recipe · Edge0 · ag-ui · awesome-codex-plugins · claude-style-patch · the whole board →