mrkeyoor.com_
Tue 29 Sept 22:50 UTC
Open Source6 min read

Strata Runs a 125B Qwen Model on One 12GB GPU

Strata turns system RAM and an SSD into working parts of local inference, putting Qwen3.8-Flash-Next within reach of a high-end gaming PC.

Strata collected 797 GitHub stars in its first four days by making a promise that usually belongs in a multi-GPU server: run a 125-billion-parameter model on one 12GB graphics card. A live check roughly three hours after that snapshot found 970 stars on the four-day-old repository. Those stars measure attention, not quality. The developer story is how Strata makes the claim work, because the graphics card is only one piece of the machine.

The project runs Qwen3.8-Flash-Next, a sparse mixture-of-experts model, through a new C++ inference engine on Windows or Linux. Qwen's official model card lists 125B parameters with 6B activated, plus a 51B n-gram embedding and a 4B multi-token-prediction layer. Strata spreads those parts across VRAM, system memory and storage, then exposes the result through local OpenAI- and Anthropic-compatible APIs.

That design widens the meaning of a local AI PC. VRAM remains the fastest and scarcest tier, but it no longer has to contain every model weight. Strata's tiered design keeps frequently used experts on the GPU, computes other experts from RAM on the CPU and reads a large lookup table from the SSD. The trade is clear: a developer can use a much larger model on one consumer GPU, provided the rest of the computer has enough memory and bandwidth to keep up.

A 12GB GPU claim with a 64GB asterisk

The minimum in the headline is real, though it is easy to misread. Strata's hardware guide recommends an RTX 30, 40 or 50 series GPU with at least 12GB of VRAM and 64GB of system RAM. The measured machine used an RTX 5070, a six-core Ryzen 5 7600 and DDR5-5200 memory. RTX 30 and 40 cards are supported by the packaged engine but were untested in the project's published table.

Storage is part of the budget too. The requirements table puts the quantized model downloads at roughly 66GB to 76GB, while the MTP layer adds about 6GB and optional image support adds another 1GB. On an AVX-512 processor, the Q2_0 configuration can create a one-time 40GB copy of its experts for a faster CPU kernel. Strata recommends an NVMe SSD, which makes sense when a 28.8GB n-gram table is being accessed during generation.

The default installation is unusually approachable for software with these requirements. On Windows, START-HERE.bat installs a private Python environment, Nvidia CUDA libraries, the engine and the selected model. Linux users run ./setup.sh. Both paths open a browser interface at 127.0.0.1:8080, and an interrupted model download can resume. The first start may stall the desktop for one to three minutes while 35GB to 55GB is loaded into RAM, a warning the project puts directly in its README.

Why sparsity changes the calculation

Qwen3.8-Flash-Next has 512 routed experts in each mixture-of-experts layer, yet each token uses only 10 routed experts and one shared expert, according to Qwen's architecture description. This lets Strata avoid treating all 125B parameters as equally urgent for every token. Across the model, the engine manages 24,576 expert blocks and learns which ones a conversation calls most often.

The GPU handles attention, the routers, shared experts, the output head and other work needed on every pass. Whatever VRAM remains becomes an adaptive expert cache. All 24,576 experts also stay pinned in system RAM, where AVX2 or AVX-512 CPU kernels compute cache misses while the GPU works on hits. The technical notes say the 28.8GB n-gram table stays on the SSD and is read a few rows at a time through the operating system's cache.

This arrangement makes PCIe traffic and memory speed part of model performance. Every extra gigabyte of VRAM holds about 700 more experts, by the project's estimate, which removes work from the CPU. Version 0.1.20 measures PCIe bandwidth at startup and reduces GPU transfers on a narrower link. A calibration command can also test the balance of CPU work, PCIe copies and speculative decoding on the user's machine.

Strata gets another speed gain from Qwen's own MTP layer. It drafts up to three tokens, then checks them together through the full model. The engine notes report an average of 2.4 to 3.2 accepted tokens per pass. A separate prompt-lookup path can reuse text already present in the context, such as code being edited, but each draft still goes through the main model before it is emitted.

The performance numbers need labels

Strata's fastest published result is 90.3 output tokens per second for the Q2_0 quantization at a 4K context. On the same RTX 5070 system, that configuration produced 67.2 tokens per second at 128K. The larger IQ3_S quantization reached 51.6 at 4K and 40.5 at 128K. These are measurements supplied by the project, using one code-agent prompt at each length and generating 256 tokens. They are useful engineering data, though they are not independent benchmarks across several PCs or workloads.

Long context also changes what speed means. The project's timing breakdown reports about four seconds to the first token for a 4K prompt with Q2_0, 25 seconds for 32K and under two minutes for 128K. At the model's native 262K limit, prompt processing takes about four and a half minutes. Once generation begins, the reply can still arrive at 60.3 tokens per second in that configuration. A single tokens-per-second figure hides that wait.

Quantization brings a quality question that star counts cannot answer. Strata offers Q2_0, IQ2_XS, IQ3_XXS and IQ3_S variants with different memory costs, and its README recommends IQ2_XS as the general choice. An optional 4-bit KV cache cut that cache's memory in half and made 128K operation about 4 percent faster in project tests, but it increased perplexity by 8 to 12 percent on long documents. The more precise 8-bit cache remains the default.

A local endpoint developers can use

The engine is more useful than a chat demo because existing tools can address it as an API. Strata's documented endpoints include OpenAI Chat Completions at /v1/chat/completions and Anthropic Messages at /v1/messages, with streaming and tool calls. A local request needs only a familiar payload:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"strata","messages":[{"role":"user","content":"Explain this patch"}]}'

The API documentation also describes a Claude Code connection through ANTHROPIC_BASE_URL. The current engine processes one request at a time and keeps one conversation's history in its live KV cache. Moving between chats can force it to read the divergent part again, although shared prefixes such as a system prompt can be reused. That makes Strata a better fit for one developer or one active agent than for a busy shared service.

By default the server listens only on 127.0.0.1. Strata can bind to a local network or sit behind a tunnel, but the network instructions tell users to set an API key first. That detail matters because the same endpoint can invoke a model, tools and MCP servers on a machine with substantial local resources. Local inference removes a cloud model call. It does not remove the need to secure a network listener.

Four days of software, moving quickly

The repository was created on September 24 and had already reached version 0.1.20 by September 28. That release added reuse of long system prompts across new chats, a PCIe-aware default and fixes for a Windows antivirus file-lock problem. The engine code is MIT-licensed, while downloaded model files retain their own licenses. The pace explains part of the interest and also argues for treating Strata as early software.

The next useful evidence will come from repeatable tests on the RTX 30 and 40 cards the project supports but has not measured, along with independent quality checks for each quantization. Reliability under long coding sessions and progress beyond the one-request queue matter more than another peak speed. If those results hold, the number to remember will still be 12GB, but only as the front door to a machine whose RAM, SSD and PCIe link are doing much of the work.

We reviewed this

  1. engine — our honest review
  2. linux — our honest review
  3. computer — our honest review

Sources

  1. Strata repository and README
  2. Strata technical details
  3. Qwen3.8-Flash-Next model card
  4. Strata v0.1.20 release notes