mrkeyoor.com_
Sat 03 Oct 13:21 UTC
Open Source6 min read

DwarfStar 4 Runs a 341GiB DeepSeek V4.1 Build on a 128GB Mac

DwarfStar 4 uses narrow model support, heavy quantization and SSD streaming to put very large open models on hardware developers can own.

DwarfStar 4 reached 246 points in the Hacker News thread that put it in MrKeyoor's October 3 brief. The number is a community-interest signal, not a performance test. The more revealing figure sits in the project's model guide: its DeepSeek V4.1 Flash Q2 file is 341 GiB, yet the software can run it on one 128 GB Mac by pulling routed experts from a fast SSD. Local inference here becomes a storage and memory-management problem well before the model fits in RAM.

The GitHub API snapshot showed 23,044 stars when checked on October 3, and the repository lists 797 commits and an MIT license. The project site identifies DwarfStar as a project by Salvatore Sanfilippo, the creator of Redis, and describes a small native inference engine for a limited set of large open-weight models. Metal is the primary target. CUDA and ROCm paths cover selected hardware and models. The repository is explicit that it is beta software and that regressions remain possible.

That narrow scope is the point. General runners try to accept many architectures and GGUF combinations. DwarfStar's repository supports project-produced GGUF files for chosen DeepSeek, GLM and Qwen families, then tests model loading, prompt rendering, tool calls, saved KV state and its HTTP server together. A model can disappear from the supported set when the maintainers find a better replacement for the machine sizes they care about.

The 341GiB file explains the design

DeepSeek V4.1 Flash Q2 contains 152 GiB of main weights plus 189 GiB of Engram tables, according to the DwarfStar model guide. Those Engram tables stay on disk and are read when needed. On a 128 GB Mac, the rest of the model also needs SSD streaming. Two 128 GB Macs or two DGX Spark systems can instead split resident Q2 inference, with each machine holding about 81 GiB of main weights while keeping a complete copy of the file on local storage.

The older DeepSeek V4 Flash Q2 path is easier to picture. Its download is about 81 GiB and is the project's starting option for 96 GB or 128 GB systems. The model guide says DwarfStar spends its most aggressive compression on the routed mixture-of-experts layers, using roughly 2-bit formats there while keeping projections, shared experts and output tensors at higher precision. The engine gains capacity by being selective about where precision is expensive, rather than applying one format to every tensor.

The same pattern extends to other model families, with materially different budgets. The supported-model table lists a 137.10 GiB Qwen3.8 Flash Next Q2 file whose main and draft weights occupy 41.73 GiB; the remaining 95.37 GiB is a disk-resident n-gram table. That is the project's 64 GB Mac starting option. GLM 5.3 Flash Q2 is about 90 GiB and targets a 128 GB Mac, DGX Spark or supported ROCm system. Model names alone do not tell you what the machine must hold.

SSD streaming buys capacity, then sends a bill

DwarfStar's SSD streaming guide is unusually direct about the compromise. Resident inference is faster. Streaming keeps a bounded cache of routed experts and fetches cache misses from the GGUF file, while ordinary weights, activations, scratch space and the context still consume memory. A model starting successfully says little about whether generation will feel interactive. Cache misses tend to hurt generation more than prompt ingestion.

The published numbers show both sides. On an M5 Max with 128 GB, a resident DeepSeek V4 Flash Q2 run produced 39.35 tokens per second at 2,048 tokens of context and 27.64 tokens per second at 65,536 tokens. The benchmark document reports 790.18 and 398.50 tokens per second for prefill at those same points. A 128 GB DGX Spark posted faster prefill in the recorded sweep, but generation ranged from 18.05 down to 13.84 tokens per second.

Those figures are upstream baselines, not independent measurements. DwarfStar's benchmark notes warn readers to hold the checkpoint, quantization, prompt and sampling settings constant when comparing runs. The benchmark restores a memory snapshot after each generation probe where possible, and its prefill rows measure only each newly added interval. A chart that mixes resident inference with streaming, or one machine with a two-node setup, answers a capacity question and a speed question at once.

Streaming examples are slower. The SSD guide reports a three-run median of 11.9 to 14.9 generated tokens per second for a 177.77 GiB GLM 5.3 Flash Q4 build on a 128 GB M5 Max. It labels a DeepSeek V4 Pro Q2 build on that machine as usable for inspection but slow. Developers should run a short generation before committing a machine to a long agent job, and should treat a warm expert cache as a different condition from a cold start.

One server can sit behind several coding agents

DwarfStar ships three ways into the same runtime: an interactive CLI, ds4-server for HTTP access and ds4-agent for persistent coding sessions. The client configuration guide documents OpenAI-compatible endpoints for Pi and OpenCode, the Responses API for Codex CLI, and an Anthropic-compatible endpoint for Claude Code. This lets a developer keep an existing agent interface while moving model execution to a machine they control.

The repository's quick start remains a source build and a large download, not a one-command desktop installer:

git clone https://github.com/antirez/ds4.git
cd ds4 && ./download_model.sh ds4f-q2
make && ./ds4

On macOS that default build selects Metal. A DGX Spark uses make cuda-spark, while a Strix Halo system has its own ROCm target. The server listens on 127.0.0.1:8000 by default, a safer default than exposing an unauthenticated compatibility API to the network. The placeholder API keys in the client examples are not authentication, a detail the client guide calls out.

Persistent KV caches are another practical part of the pitch. The repository documentation says the native agent can save sessions under ~/.ds4/kvcache, resume compatible snapshots and avoid rebuilding an unchanged prompt prefix. Stripping a session keeps its text while removing the large KV payload. Images currently prevent a native-agent session from being saved, and network tensor-parallel restores still require prefill. Local storage can therefore contain conversation traces as well as model weights, so it belongs in the same privacy review as any cloud log.

Beta means reading the matrix before downloading

The repository's 23,000-plus stars do not widen its compatibility matrix. The model guide says DeepSeek V4.1 vision currently requires Metal. Its CUDA text path supports selected Q2 arrangements, while ROCm and pipeline execution are not implemented for that model. Qwen3.8 Flash Next and GLM 5.3 have different file sizes, runtime allocations and backend coverage. An arbitrary GGUF with a familiar model name may still have an unsupported layout or quantization mix.

Testing is present, but the maintainers define its boundaries. The testing guide separates model-free checks, small GPU tensor tests and checkpoint-matched output tests. It also says a smoke test is not full quality assurance. That distinction matters for an inference engine built with what the repository calls strong assistance from coding agents: integration evidence has to follow the exact model, backend and execution mode being shipped.

The next useful evidence will come from repeatable results outside the project's own hardware, especially cold-cache measurements and long coding sessions on 64 GB and 128 GB machines. The benchmark guide already asks testers to separate warm and cold caches and report individual latency beside aggregate throughput. Watch whether contributors can keep DeepSeek V4.1, GLM and Qwen support current without turning the narrow engine into another general runner. DwarfStar's bet is clearest when a 341 GiB file runs on a machine with less than half that memory. Its credibility will depend on how predictable that mismatch becomes.

We reviewed this

  1. ds4 — our honest review

Sources

  1. DwarfStar 4 project site
  2. antirez/ds4 repository
  3. GitHub API metadata for antirez/ds4
  4. DwarfStar model and hardware guide
  5. DwarfStar performance and benchmarking guide
  6. DwarfStar SSD streaming guide
  7. DwarfStar coding agent client guide
  8. DwarfStar testing guide