mrkeyoor.com_
Tue 06 Oct 17:49 UTC
AI6 min read

Mistral Large 4 Activates 49B Parameters Across a 1.05T Model

Mistral's sparse design cuts active compute to 49B parameters, but its promised 1.05T-parameter checkpoint leaves a much larger self-hosting problem.

A Hacker News submission linking to Mistral Large 4's model card collected 821 points and 456 comments in the brief's 13:30 UTC snapshot, but the number with a longer shelf life is 21.4. The model has 1.05 trillion parameters and activates 49 billion of them during inference, a ratio of roughly 21.4 to one. That gap explains both the appeal and the catch: sparse computation avoids applying all trillion parameters to each token, while the promised checkpoint still leaves operators with a trillion parameters to store, move and distribute across accelerators. The Hacker News discussion shows how quickly developers noticed the release; it does not answer the deployment questions.

Mistral launched ML4 on October 6 as a public API preview. The company says the weights will arrive by the end of October, alongside more architecture detail, extra evaluations and its post-training method. Until then, developers can test the hosted model, but they cannot inspect the checkpoint or reproduce the serving setup. Mistral's announcement says the model was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in the company's European data centers and that the preview is served from its own infrastructure. It does not describe how many GPUs handle one inference request.

The 49B figure describes activity, not model size

ML4 uses what Mistral calls a granular mixture-of-experts architecture. Instead of applying every parameter to every input, the model routes work through a smaller active portion. The official model card lists 49 billion active parameters, 1.05 trillion total parameters and a separate 1.6 billion-parameter vision encoder. Mistral has not yet published the routing scheme, expert count or the number of experts selected for a token, so the 49B figure is useful as a compute signal rather than a complete performance specification.

The storage arithmetic is less forgiving. At 16 bits per parameter, 1.05 trillion parameter values alone work out to about 2.1 terabytes. An 8-bit representation is roughly 1.05TB, while a 4-bit representation is about 525GB. Those are lower-bound calculations from the published parameter count, not measured ML4 package sizes. Checkpoint metadata and runtime buffers add overhead. The encoder adds more storage if Mistral's 1.05T count excludes it. Some quantization methods also cost accuracy, and Mistral has not supplied local quality results for any compressed build.

That distinction matters when a team hears 49B and starts pricing a server as if ML4 were an ordinary 49B dense model. Sparse routing can reduce the work performed for a token, but every expert needed by the model still has to be available somewhere in the serving system. Operators may shard those weights across machines or move experts between storage tiers, with latency and network costs determined by an architecture that Mistral has not disclosed. The announcement's end-of-month promise leaves hardware planning premature today.

The one-million-token context window on the model card adds another variable. Long-context inference needs memory beyond the checkpoint for the attention cache and runtime state. The amount depends on implementation details, sequence length, concurrency and precision. A model may fit after aggressive weight compression yet still run out of room when several long jobs arrive together. The public 1M limit tells API users what they can request; it does not provide a memory budget for a local deployment.

An API preview is available before the open weights

Mistral calls ML4 open-weight, but the two words describe a release that has not happened yet. At publication time, the model card labels version 26.10 as a public preview and lists the model identifier mistral-large-4. Its Weights pane does not contain a download, checksum or license. Neither the card nor the launch post names the license that will govern the files.

There is still a useful product to test. The model card lists chat completions, structured output, function calling, document question answering, batching and Mistral's Agents and Conversations endpoints. It also advertises multimodal input and a one-million-token context window. Those interfaces let teams run their own documents, tool schemas and coding tasks against the preview instead of relying on launch charts. They cannot reveal checkpoint size, local throughput or the operational behavior of an eventual quantized build.

Security evaluators face a separate access question. Mistral says selected cybersecurity leaders, vetted partners and state authorities are red-teaming the same model with reduced moderation and expanded cyber capabilities before the weight release. The public announcement does not offer that access level to every API customer. A team testing the ordinary preview should therefore record refusals alongside task success instead of assuming it is seeing the policy configuration used for Mistral's cyber results. The company's description of the red-team program also gives the weight release a second job: it must show what safeguards travel with a self-hosted deployment and which controls remain part of Mistral's hosted service.

The checkpoint may also land after the model has changed. Mistral says the reinforcement-learning run behind the preview remains in flight and produces about 33 billion tokens per day on 3,000 GPUs, with around 16 billion retained as trainable completion tokens after filtering and masking. The company expects the model to keep improving before the weights ship. That is an unusually direct warning against treating today's API result as a frozen artifact. Mistral has not said whether the October checkpoint will match the current preview exactly.

The benchmark claims need their labels

Mistral reports a 61.7 percent score on DeepSWE v1.1 and 28.3 percent on Terminal-Bench 4, producing a 49.8 percent score on its combined Coding Agent Index. In a blind coding evaluation run with Surge AI, the launch post gives ML4 a mean rating of 3.74 out of five, behind Claude Opus 5 at 4.22 and ahead of the three other models shown. These are more informative than a single general leaderboard because they separate repository work, terminal use and judged output quality. They are still results selected and presented by Mistral.

Cybersecurity is the bolder part of the pitch. Mistral says ML4 scored 82 percent on a test that requires reproducing and patching real open-source vulnerabilities, and 93 percent on the 40-challenge Cybench set. The same post reports 93.3 percent resistance on Lakera's public B3 indirect-prompt-injection benchmark. Mistral frames the combination as a way for defenders to run capable security models under their own policies, including in private cloud or on premises. That on-premises claim becomes testable only when the weights, safeguards and serving guidance are public.

The company also says ML4 was trained with data spanning more than 160 languages and can inspect documents, charts and natural images. Its strongest advertised uses extend beyond chat into technical drawings, geospatial imagery, spreadsheets and legal or financial work. The model card's 1.6B vision encoder confirms that vision is part of the architecture, though it provides no separate resource requirement or evaluation method for that encoder. Buyers should test the exact modality they need rather than transfer a coding score to a document or vision workflow.

The files will settle the deployment question

The end-of-month release needs to answer several practical questions left open by the preview. A weight manifest will show the real download and shard sizes. The license will define what companies can modify and redistribute. Architecture documentation should explain expert routing, supported precisions and the reference serving stack. Reproducible evaluations would let outside teams check whether the downloadable checkpoint behaves like the hosted model described in the October 6 announcement.

For now, the sensible test is split in two. API users can measure quality, latency and tool behavior with their own workloads through mistral-large-4. Self-hosters should hold off on accelerator counts until the checkpoint and routing details exist. If Mistral meets its October deadline, the 21.4-to-one ratio will move from an interesting model-card number to a concrete systems problem, with file sizes, licenses and local benchmarks that anyone can inspect.

We reviewed this

  1. terminal — our honest review
  2. runtime — our honest review
  3. Files — our honest review

Sources

  1. Mistral Large 4 announcement
  2. Mistral Large 4 model card
  3. Hacker News discussion: Mistral Large 4