mrkeyoor.com_
Sat 03 Oct 17:29 UTC
AI6 min read

Kolibri 1 Pairs Apache-2.0 Weights With a 78GB Memory Footprint

Aleph Alpha's bilingual mixture-of-experts model can run on private infrastructure. Its 3.46B active parameters hide a 78GB memory footprint and new vLLM plumbing.

At MrKeyoor's briefing snapshot, Kolibri's company-release listing had 192 Hacker News points and 30 comments. A second listing for an independent explainer had 297 points and 124 comments. The more useful number for developers is 3.46 billion: that is how many parameters work on each token, though the full model still occupies about 78GB of GPU memory. Kolibri is computationally sparse and physically large.

Aleph Alpha released Kolibri 1 on October 3 as an English-German mixture-of-experts model with downloadable weights. The company's launch post lists an Apache 2.0 license, a context window of up to 1,048,576 tokens, tool calling and four reasoning settings. The release is aimed at organizations that want to keep documents and inference inside infrastructure they control.

The 3.46B label needs a memory footnote

Kolibri has 78.1 billion parameters spread across 50 mixture-of-experts layers. Each layer contains 384 routed experts, and six are selected for a token alongside one shared expert. That routing cuts the computation per token to 3.46 billion active parameters, or 4.4 percent of the full model, according to Aleph Alpha's 189-page technical report.

The inactive experts do not disappear from the machine. Kolibri's FP8 weights occupy about 78GB, and the model card gives a minimum of two 80GB A100s, two H100 SXM5 GPUs, or one H200, B200 or B300. That hardware list leaves capacity above the weight footprint for the runtime and key-value cache. A developer reading only the active-parameter count could reasonably expect laptop-class hardware. This model belongs on a server.

That trade has a purpose. A sparse model can draw on the capacity of many experts without applying every parameter to every token. Aleph Alpha's architecture tests found that a larger 123B design handled three concurrent 256,000-token requests on two H100s, while the chosen 78B design handled 18 and decoded 28 percent faster. Those are company measurements from a specific test setup, so they should not be read as a promise for every deployment.

Long context adds another practical qualification. Kolibri was trained at 16,384 tokens, extended through later stages to 262,144, then tested at 1,048,576. The model card recommends staying at or below 262,144 tokens for complex work or when latency and throughput matter. One million tokens is an available ceiling rather than the routine serving target.

What sovereign means here

Aleph Alpha uses sovereignty to describe both the production chain and the customer's control after delivery. Its teams built the model in Germany and trained it on infrastructure in Germany and Finland, under German and European law, according to the release account. A customer can download the weights and run inference without sending private prompts to an outside API.

The word does not mean every ingredient originated in Europe. The training disclosures say the data pipeline used outside models for some synthetic text and quality filtering. The model card also warns that Chinese models used in data preparation can carry political bias, which Aleph Alpha says it addressed through filtering and alignment work. Tejas Kumar's technical reading of the release makes the useful distinction: deployment control can be local even when the model's upstream dependencies are international.

Kolibri is open-weight, a narrower description than open source. Aleph Alpha's license note says Apache 2.0 applies to the weights and configuration files in the Hugging Face repository. It expressly excludes other artifacts and says the company retains rights to its code, model architecture, training methods and parameter settings. Teams can inspect, run and adapt the released weights. They do not receive the complete training system required to reproduce Kolibri from raw data.

The Hugging Face repository reports about 78.85GB of stored files, including 32 weight shards, configuration files and a tokenizer. Aleph Alpha also published the technical report and a public training-data summary. Evaluators get more material than a hosted API alone, while reproducibility stops short of a source release for the full pipeline.

German is part of the architecture

Nearly a quarter of Kolibri's 20 trillion pre-training tokens are German, with English at about 62.5 percent and code at 13.6 percent. Mid-training added 3.44 trillion tokens, followed by 201 billion tokens for long-context extension. Aleph Alpha trained the model on 768 Nvidia B200 GPUs; the model card reports 511 hours and 392,000 GPU-hours for pre-training, excluding the later stages.

The same card estimates 950MWh of energy for pre-training, mid-training and long-context work, including data-center overhead. It excludes supervised fine-tuning, reinforcement learning, ablation models and low-load states. That estimate concerns model creation rather than inference, yet it adds useful scale to Aleph Alpha's claim of controlling the production chain.

The bilingual focus reaches the tokenizer. German compounds often split into more tokens when a tokenizer was built mainly around English. Kumar tested Kolibri's tokenizer on the German Basic Law and counted 35,190 tokens, compared with 41,482 for OpenAI's o200k_base tokenizer. It is one independent experiment on one legal document. The result shows why tokenization affects cost and usable context for the audience Kolibri targets.

Aleph Alpha's evaluation tables put Kolibri at 75.5 on an aggregate English suite and 70.8 on German among the post-trained models tested. Qwen3.5 35B-A3B scored 74.7 and 69.8 in the same table, while the dense Qwen3.8 27B scored higher at 80.2 and 79.9. The comparison supports a narrower claim than "best German model": Kolibri tested well against similarly sparse models while using fewer active parameters than dense rivals. Independent replications will matter because the developer also chose the tasks, prompts and serving setup.

The report measures throughput in decoded bytes per GPU because tokenizers divide text differently. Its headline comparison ran each model on eight B200 GPUs, with 4,000-token sequences for base models and 16,000 for post-trained models. That setup makes the models comparable inside Aleph Alpha's test. It does not estimate performance on the minimum two-GPU configuration listed for Kolibri.

The weak rows are as informative as the aggregate. Kumar's review of the published tables found Kolibri behind Qwen models on Terminal-Bench 2.1 and SWE-bench Verified, and weaker on multi-turn function calling. Aleph Alpha positions it for document work, retrieval and German-language reasoning. A coding-agent team should test those lower scores before treating the broad average as a buying guide.

Running it requires Aleph Alpha's vLLM plugin

The first public build is not a drop-in model for an unmodified vLLM install. Aleph Alpha provides an aleph-alpha-inference package that installs the supported vLLM version and registers Kolibri's reasoning and tool-call parsers. The model card's serving command is short:

pip install 'aleph-alpha-inference>=1'

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

That starts an OpenAI-compatible endpoint. Applications can set reasoning effort to none, low, medium or high through the chat template. Reaching the advertised million-token ceiling requires two more serving flags, and memory use rises with the cache. The 78GB weight footprint is only the starting figure for capacity planning.

The model card's guidance also draws a firm operational boundary around agents. It recommends human review before actions and says tool results should be validated by the calling system. For sensitive decisions, Kolibri is meant to advise a person rather than make the decision itself. Open weights move inference under the operator's control. They do not supply application policy, output filters or safe tool permissions.

The next evidence should come from operators

The two Hacker News discussions show strong launch-day curiosity, though neither thread verifies the vendor's performance claims. The useful next results will be reproducible serving numbers on H200 and B200 hardware, evaluations on private German documents, and reports on how the new vLLM plugin behaves under sustained load.

Kolibri gives regulated teams something concrete to test: downloadable bilingual weights, detailed disclosures and an on-premise route with no inference API in the middle. The test begins with a less glamorous number than its benchmark score. Budget for the full 78GB model, the cache and the integration work, then measure whether control over the deployment is worth that hardware bill.

We reviewed this

  1. Files — our honest review
  2. pipeline — our honest review
  3. tokenizers — our honest review

Sources

  1. Kolibri Has Landed: A Sovereign Open-Weight Model
  2. Kolibri 1 model card
  3. Kolibri: A Sovereign European Model on the Pareto Frontier
  4. Aleph Alpha Kolibri: How the Sovereign German LLM Works
  5. Kolibri launch discussion on Hacker News
  6. Kolibri technical analysis discussion on Hacker News