mrkeyoor.com_
Fri 02 Oct 15:00 UTC
LLM Toolsevaluationupdated 26 Aug 2026

colibri review

Colibri is a compact C inference engine that streams mixture-of-experts model weights across storage, system memory, and optional GPU memory. It is built for running unusually large open-weight models on hardware that cannot keep every expert in fast memory, with chat, server, web, and planning commands around the engine.

+1,096stars / 7d
Verdict

Our Colibri run installed 34 Python packages, built in 10 seconds, and found 0 known vulnerabilities, but it ran no tests and loaded no model weights. Colibri is worth studying if expert streaming itself is the problem you want to solve and you can dedicate hundreds of gigabytes of storage to the experiment. For ordinary local inference or a supported service, start with llama.cpp, KTransformers, or vLLM and move only when Colibri's model roster and memory hierarchy answer a measured constraint.

We ran it

Lab card: what happened when we ran colibriScreenshot of colibri (justvugg.github.io/colibri)
Install✓ · 21s34 packages · 36 MB
Build✓ · 10s
Testsn/ano test script
Known vulns0(pip-audit)
Repo495 files~108,867 lines of source · 13.7 MB · 4 CI workflows

Answers from our run

Does colibri build from source?

Dependencies installed in 21 seconds (34 packages), and the build succeeded in 10 seconds. We cloned commit 33e67a9 into a clean Debian container with 3 CPUs and no project-specific setup.

Does colibri have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does colibri have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use colibri?

Users expecting a quick local chatbot download: the README's reference GLM container is about 372 GB, while Kimi K3 is about 1.6 TB.

What are the alternatives to colibri?

llama.cpp, KTransformers, vLLM. Our Colibri run installed 34 Python packages, built in 10 seconds, and found 0 known vulnerabilities, but it ran no tests and loaded no model weights.

Setup2/5Small launcher build, but useful models need major storage and tuning
Docs5/5Detailed model, backend, tuning, experiment, and hardware guidance
Community5/526,225 stars with same-day August 2026 issues and pull requests
Maturity3/5Six model engines work, while several hardware paths remain experimental

Discussed on

  1. hnShow HN: Getting GLM 5.2 running on my slow computer937 points
  2. hnShow HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri48 points

Who it’s for

Inference engineers studying expert caching, storage I/O, CPU and GPU overlap, and model-specific kernels.
Advanced local-model users with large, fast storage and patience for hardware-specific tuning.
Researchers willing to publish controlled measurements, including negative results.
Teams that can validate model containers, output quality, resource plans, and failure behavior themselves.

Who it’s NOT for

Users expecting a quick local chatbot download: the README's reference GLM container is about 372 GB, while Kimi K3 is about 1.6 TB.
Teams that need predictable latency or a support SLA: the README explicitly provides no speed SLA and labels several GPU paths experimental.
Operators who want one engine for arbitrary transformer checkpoints: Colibri has separate C engines for six named model families and model-specific containers.
Buyers who require a measured runtime test before adoption: our Python-oriented run built successfully but exposed no test target and never loaded a model.
Windows users unwilling to troubleshoot release artifacts: open issue 1241 reports current Windows 11 executables flashing a terminal and exiting without a useful log.

Setup reality

Our Python sandbox installed 34 packages in 21 seconds and used 36 MB. The build succeeded in 10 seconds, and pip-audit reported 0 known vulnerabilities. No test script or target was available, so tests were skipped. The checkout had 495 files, about 108,867 source lines, and occupied 13.7 MB.

That small package result does not include model weights. Prebuilt releases need Python 3 for the launcher and gateway; source builds need GCC or Clang with OpenMP. The engine selected for a model may add CUDA, Metal, Vulkan, storage, or platform requirements.

Model preparation is the real setup. The documented families range from an approximately 7 GB OLMoE container to a roughly 1.6 TB Kimi K3 checkpoint. Colibri offers planning and doctor commands, but storage capacity, RAM headroom, model format, and backend compatibility still need a machine-specific check.

Colibri treats model weights as a storage hierarchy

Colibri targets mixture-of-experts models whose full parameter count is much larger than the subset used for one token. Dense weights stay resident while routed experts can live in VRAM, RAM, or storage. A per-layer cache, recorded routing history, and lookahead decide which experts move closer to compute. The stated rule is that placement may change speed but must not silently change model precision or routing semantics.

That approach is specific to sparse model architecture. It does not make a dense model smaller, and it does not remove the bytes needed for the checkpoint. The reference GLM-5.2 container is about 372 GB. Inkling is about 469 GB, while Kimi K3 is roughly 1.6 TB. OLMoE and Qwen3.6 offer smaller entry points, but every family has its own model format, engine, memory plan, and backend notes.

Six model families mean six model-specific engines

The README lists GLM-5.2, Inkling, Kimi K3, DeepSeek V4 Flash, Qwen3.6, and OLMoE. Each family has a separate C file over shared headers for tokenization, safetensors access, quantization, expert storage, routing traces, and prefix reuse. The common coli chat, coli serve, and coli web commands inspect model configuration and select the matching binary.

This design keeps architecture details visible, which suits research and debugging. It also sets a firm compatibility boundary. A new MoE checkpoint does not become supported merely because its weights use safetensors. Attention, routing, expert layout, tokenizer, chat template, quantization, and output semantics must match an engine. Teams should verify the exact container named by the documentation instead of substituting a similarly labeled quantization.

What happened when we ran it

Our sandbox cloned commit 33e67a9 into a fresh unprivileged Debian container with 3 CPUs and 8 GB of RAM. The checkout contained 495 files, about 108,867 lines of source, and occupied 13.7 MB. The detected Python project installed 34 packages in 21 seconds and used 36 MB on disk. Pip-audit reported 0 known vulnerabilities in that installed environment.

The build succeeded in 10 seconds. No test script or target was available, so the harness skipped tests. The repository had 4 CI workflow files, no Dockerfile, and no tests directory. The README's layout describes dependency-free C and Python checks under c/tests, but the generic measured path did not discover a runnable test target. We therefore have no passing test count from this commit.

Our run did not compile every C engine, download a checkpoint, validate tensors, start the API, generate a token, exercise a GPU backend, or measure storage throughput. The 36 MB installed footprint belongs to the Python-side project, not a usable model deployment. Any claim about inference speed or quality would require a separate run on declared hardware with a specific model and container.

Planning commands help, but they cannot create headroom

Prebuilt releases include a Python launcher around the C engine. Source builds need GCC or Clang with OpenMP, while optional paths add CUDA, Metal, or Vulkan. coli plan shows the proposed placement, coli doctor checks readiness, and coli tune records a machine profile. These are useful before committing to a long model conversion or startup.

Capacity remains unforgiving. Model files need enough disk space, conversion may need temporary space, and resident dense weights need real RAM. Extra cache can push a large process near the operating system's limit. Open issue 766 describes an older auto-tier plan whose chosen cap drove a 251 GiB host into the Linux OOM killer. Version 1.8.0 says tuning now measures safe RAM and cache caps, but operators should still leave headroom and test forced failure rather than trusting one projection.

Hardware backends need proof on the exact machine

Colibri can combine CPU execution with optional CUDA, Metal, or Vulkan expert tiers. It also supports multiple SSDs, NUMA placement, prefetch, compressed cache state, and experimental speculative paths. Many switches are intentionally opt-in because a faster micro-operation may lose across a whole decoding turn. That experimental discipline is good, though it means a copied environment-variable recipe may be slower or less stable on another host.

Open issue 813 reports a Metal build announcing GPU expert work while its counters showed none on an M3 Ultra. The report is tied to a development commit and unusual full-residency hardware, so it should not be generalized to every Apple system. It does show why backend banners are insufficient: check counters, output identity, memory use, and an end-to-end result before believing that a GPU path is active.

Release activity is intense for a young research engine

Version 1.8.0 was published on August 24, 2026 after 60 pull requests since v1.7.0. Its notes cover constant-time expert lookup, prefill caching, older CUDA architectures, experimental hybrid execution, Kimi checkpoints, KV-cache formats, tuning, platform fixes, CI, and documentation. GitHub showed 26,225 stars, a push on August 25, and 110 combined open issues and pull requests.

Colibri is unusually candid about hypotheses, negative results, model-specific tradeoffs, and measurement procedure. That makes it valuable for inference engineers. It does not make a 372 GB checkpoint convenient, nor does our 10-second Python build prove the C runtime. Adopt it as a measured systems project with pinned weights and raw logs, not as a drop-in local chat application.

Alternatives

ProjectWhat it isPick it when
llama.cpp gh↗A widely used C and C++ inference runtime covering many model families and hardware backends.pick this instead when broad model support, established quantized formats, and a larger user base matter more than Colibri's expert-streaming experiments.
KTransformersAn inference system focused on flexible CPU and GPU execution for large transformer models.pick this instead when its supported architectures and heterogeneous execution path match the model and hardware you own.
vLLM gh↗A production-oriented inference server built for throughput and supported accelerator deployments.pick this instead when serving concurrent users on suitable GPUs is the job, rather than streaming frontier MoE experts from disk.

What people are saying

  1. [velocity-scout] JustVugg/colibri

Sources

  1. Colibri README
  2. Colibri v1.8.0 release
  3. Colibri issue 766
  4. Colibri issue 813
  5. Colibri issue 1241

More llm tools reviews

whatsapp-mcp · deepseek-recipe · Edge0 · ag-ui · awesome-codex-plugins · claude-style-patch · the whole board →