mrkeyoor.com_
Mon 05 Oct 04:09 UTC
Open Source7 min read

Strata's 610-Point Qwen Run Leaves a Two-Bit Quality Gap

Strata crossed 11,000 GitHub stars after users reported 100-plus tokens/s on RTX 4090 PCs. The fastest Q2 path still lacks independent task testing.

A 610-point Hacker News thread pushed Strata past 11,000 GitHub stars in its eleventh day, powered by RTX 4090 users reporting more than 100 tokens per second from a 125-billion-parameter Qwen model. That response changes the question around this open-source inference engine. Strata has already shown that Qwen3.8-Flash-Next can run quickly across a gaming GPU, system RAM, and an SSD. Now developers need to know what its fastest two-bit build loses on real work.

The Hacker News submission reached 610 points and 282 comments in under 13 hours. Its author reported 124 tokens per second on an RTX 4090, a Ryzen 9 7950X3D, and 128GB of DDR5 memory. Another 4090 owner in the thread reported more than 110 tokens per second with a 60,000-token context. Those accounts are useful field reports, not controlled tests. They made the speed claim believable enough for thousands of developers to look closer.

Strata's repository had 11,118 stars and 977 forks when checked on October 5. It was created on September 24 and had already reached version 0.1.39. Independent task evaluations have not yet caught up with the current release.

The 100-token result has identifiable conditions

Strata's own numbers support the general shape of the user reports. Its measured benchmark uses an RTX 5070 with 12GB of VRAM, a Ryzen 5 7600, and 64GB of DDR5-5200. With the Q2_0 quant and a short context, engine 0.1.36 produced 93.5 tokens per second. At a 128,000-token context, it produced 76.4.

The project estimates that an RTX 3090 with 24GB of VRAM will produce about 100 to 140 tokens per second at short context, with a stated error band of plus or minus 20 percent. An RTX 4090 has the same VRAM capacity but much higher memory bandwidth, so the two community reports are plausible. They are not a guarantee. CPU throughput, RAM speed, PCIe bandwidth, quant choice, prompt length, and the text being generated all affect Strata's result.

Part of the speed comes from multi-token prediction. A smaller draft layer proposes several tokens, then the large model verifies the group. Strata says this makes output 1.6 to 1.8 times faster. Acceptance changes with the answer, however, so two prompts on the same machine can produce different rates. A user in the thread reporting 110-plus tokens per second specified a three-token MTP setting, one of the details a headline leaves out.

Prompt processing also has its own clock. Strata measured 2,653 tokens per second while ingesting a 32,000-token prompt with Q2_0 on the RTX 5070 machine. Output speed describes what happens after that context is read. For a coding agent that repeatedly sends a large repository view, prompt delay and cache reuse may matter more than how quickly the final response streams.

Q2 is doing much of the headline work

Qwen3.8-Flash-Next is unusually suitable for this kind of split execution. Qwen's model card lists 125B parameters with 6B activated for each token, plus a 51B n-gram embedding and a 4B multi-token-prediction component. Each of its 48 layers has 512 routed experts, while a token selects 10 plus one shared expert. Strata keeps frequently used experts in VRAM, computes other experts from system memory, and leaves the large n-gram table on the SSD.

The fast configuration compresses those weights to roughly two bits. Strata's model guide calls Q2_0 its fastest option and assigns it the broad quality label "good." IQ2_XS uses slightly more memory, runs at 79 tokens per second in the project's short-context RTX 5070 test, and gets the label "better." IQ3_S reaches 53 tokens per second and is described as the best native option.

Those labels help someone choose an install, but they do not say how often a coding agent fixes the right bug, preserves a constraint after 20 tool calls, or returns valid structured output. The project's speed tables are unusually detailed. Its public quality comparison between the quant levels is much thinner. That mismatch matters because the HN headline joins three facts that readers may treat as one: 125B parameters, 100 tokens per second, and the capability of the original model.

Hardware requirements narrow the meaning of "consumer" too. Q2_0 needs an estimated 37.6GB across RAM and VRAM. Strata recommends at least 12GB of VRAM, 32GB of system RAM for its reduced Coder variant, and 48GB for Q2_0 or IQ2_XS on most configurations. Downloads for the smaller full-expert versions run from 66GB to 76GB, with another roughly 6GB for the draft layer. The 4090 reports used 128GB of system memory.

Qwen's benchmark table cannot fill the gap

Qwen publishes a wide set of results for the original Qwen3.8-Flash-Next weights. Its model card reports 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual, and 48.1 on NL2Repo-Bench under the listed harnesses. Those are Qwen's own evaluations, and several use in-house or modified setups. More important here, they do not measure Strata's Q2_0, IQ2_XS, or IQ3_S packages.

Quantization can preserve much of a model's average benchmark score while changing failures that agents expose. A rare tool token, a JSON delimiter, or a low-probability instruction can matter more than a small aggregate score movement. Strata's technical notes acknowledge related precision effects. Its optional four-bit KV cache cuts cache memory in half and speeds a 128K run by about 4 percent, while raising perplexity by 8 to 12 percent on long documents. The eight-bit KV cache remains the default.

The reduced Coder build offers one concrete comparison, though it answers a different question. It removes half of each layer's experts, selected using code data, and fits a 32GB PC. Its authors report 91 percent of the full model's SWE-bench Verified result and 99 percent of its LiveCodeBench result. Strata also warns that the build is weaker outside code and has produced wrong or looping Chinese output in a reported issue. That is the kind of task-level limit the Q2 variants need documented.

The community evidence is early and uneven

Several Hacker News commenters said Strata worked well for coding, including one who mentioned PHP security audits. Others reported around 60 tokens per second with IQ3_S on an RTX 3090 and 65 with IQ2 on a Radeon RX 9070 XT. These accounts widen the hardware sample, but they lack shared prompts, output caps, warm-up rules, or correctness checks. They show that people can run the engine. They cannot rank its quants.

One commenter supplied a more pointed result: a 50-image coordinate test in which Strata had a median error of 154.8 pixels, compared with 46.5 for the same GGUF and vision adapter through llama.cpp. The author reported temperature zero but did not publish a complete test package in the thread. It is a lead worth reproducing, not a settled verdict on Strata's vision path. The thread also contains positive comparisons with Qwen 27B and skeptical reports about accuracy. Community buzz is exposing the right questions before it has produced dependable answers.

Strata's community benchmark template asks contributors to record prompts, repetitions, cache state, loading, timing boundaries, and memory. Its results table is still a blank form. Filling that page with linked logs would be more useful than collecting isolated tokens-per-second screenshots. A good comparison would run the same coding tasks through Q2_0, IQ2_XS, IQ3_S, and a smaller model at higher precision, then score both completion and tool-use failures.

Version 0.1.39 broadens the operational test

Strata 0.1.39, published shortly before the HN submission, adds the OpenAI Responses API used by Codex, opt-in parallel requests, faster decode kernels, and more hardware paths. The release says its decode changes improved Q2_0 by 6 percent on two RTX 5070 workloads, using 10 interleaved pairs against version 0.1.38. It also reports byte-identical output in those fixed-cache comparisons.

Parallel service brings another useful trade. With four requests on the RTX 5070, the last request began after 1.8 seconds instead of 11.2, but total decoding was 11 percent slower. A request running alone lost 11 percent with two slots and 22 percent with four because each slot takes 0.56GB from the expert cache. Strata remains better suited to one developer or a small number of active agents than to a busy shared endpoint on a 12GB card.

The server listens on 127.0.0.1 by default and supports local OpenAI-style and Anthropic-style APIs. Its security documentation says to set an API key before binding to 0.0.0.0 or placing a tunnel in front of it. It also states that the project has not had an outside security audit. MCP tools are opt-in and run with the user's permissions, a detail that deserves the same care as model accuracy when an agent can edit files or execute commands.

One result would now tell developers more than a higher decode peak: a reproducible success rate for the exact Q2 build behind Strata's 100-token claim, measured on long coding and tool-use tasks beside IQ3_S and a smaller, less compressed model. Until that table exists, 610 Hacker News points measure developer appetite more clearly than they measure the model developers will get.

We reviewed this

  1. table — our honest review
  2. engine — our honest review
  3. requests — our honest review

Sources

  1. Strata repository
  2. Strata technical details and benchmarks
  3. Strata model guide
  4. Strata community benchmark template
  5. Qwen3.8-Flash-Next model card
  6. Hacker News discussion: Run Qwen 3.8 Flash Next on consumer hardware
  7. Strata v0.1.39 release notes
  8. Strata security documentation