Colibri treats model weights as a storage hierarchy
Colibri targets mixture-of-experts models whose full parameter count is much larger than the subset used for one token. Dense weights stay resident while routed experts can live in VRAM, RAM, or storage. A per-layer cache, recorded routing history, and lookahead decide which experts move closer to compute. The stated rule is that placement may change speed but must not silently change model precision or routing semantics.
That approach is specific to sparse model architecture. It does not make a dense model smaller, and it does not remove the bytes needed for the checkpoint. The reference GLM-5.2 container is about 372 GB. Inkling is about 469 GB, while Kimi K3 is roughly 1.6 TB. OLMoE and Qwen3.6 offer smaller entry points, but every family has its own model format, engine, memory plan, and backend notes.
Six model families mean six model-specific engines
The README lists GLM-5.2, Inkling, Kimi K3, DeepSeek V4 Flash, Qwen3.6, and OLMoE. Each family has a separate C file over shared headers for tokenization, safetensors access, quantization, expert storage, routing traces, and prefix reuse. The common coli chat, coli serve, and coli web commands inspect model configuration and select the matching binary.
This design keeps architecture details visible, which suits research and debugging. It also sets a firm compatibility boundary. A new MoE checkpoint does not become supported merely because its weights use safetensors. Attention, routing, expert layout, tokenizer, chat template, quantization, and output semantics must match an engine. Teams should verify the exact container named by the documentation instead of substituting a similarly labeled quantization.
What happened when we ran it
Our sandbox cloned commit 33e67a9 into a fresh unprivileged Debian container with 3 CPUs and 8 GB of RAM. The checkout contained 495 files, about 108,867 lines of source, and occupied 13.7 MB. The detected Python project installed 34 packages in 21 seconds and used 36 MB on disk. Pip-audit reported 0 known vulnerabilities in that installed environment.
The build succeeded in 10 seconds. No test script or target was available, so the harness skipped tests. The repository had 4 CI workflow files, no Dockerfile, and no tests directory. The README's layout describes dependency-free C and Python checks under c/tests, but the generic measured path did not discover a runnable test target. We therefore have no passing test count from this commit.
Our run did not compile every C engine, download a checkpoint, validate tensors, start the API, generate a token, exercise a GPU backend, or measure storage throughput. The 36 MB installed footprint belongs to the Python-side project, not a usable model deployment. Any claim about inference speed or quality would require a separate run on declared hardware with a specific model and container.
Planning commands help, but they cannot create headroom
Prebuilt releases include a Python launcher around the C engine. Source builds need GCC or Clang with OpenMP, while optional paths add CUDA, Metal, or Vulkan. coli plan shows the proposed placement, coli doctor checks readiness, and coli tune records a machine profile. These are useful before committing to a long model conversion or startup.
Capacity remains unforgiving. Model files need enough disk space, conversion may need temporary space, and resident dense weights need real RAM. Extra cache can push a large process near the operating system's limit. Open issue 766 describes an older auto-tier plan whose chosen cap drove a 251 GiB host into the Linux OOM killer. Version 1.8.0 says tuning now measures safe RAM and cache caps, but operators should still leave headroom and test forced failure rather than trusting one projection.
Hardware backends need proof on the exact machine
Colibri can combine CPU execution with optional CUDA, Metal, or Vulkan expert tiers. It also supports multiple SSDs, NUMA placement, prefetch, compressed cache state, and experimental speculative paths. Many switches are intentionally opt-in because a faster micro-operation may lose across a whole decoding turn. That experimental discipline is good, though it means a copied environment-variable recipe may be slower or less stable on another host.
Open issue 813 reports a Metal build announcing GPU expert work while its counters showed none on an M3 Ultra. The report is tied to a development commit and unusual full-residency hardware, so it should not be generalized to every Apple system. It does show why backend banners are insufficient: check counters, output identity, memory use, and an end-to-end result before believing that a GPU path is active.
Release activity is intense for a young research engine
Version 1.8.0 was published on August 24, 2026 after 60 pull requests since v1.7.0. Its notes cover constant-time expert lookup, prefill caching, older CUDA architectures, experimental hybrid execution, Kimi checkpoints, KV-cache formats, tuning, platform fixes, CI, and documentation. GitHub showed 26,225 stars, a push on August 25, and 110 combined open issues and pull requests.
Colibri is unusually candid about hypotheses, negative results, model-specific tradeoffs, and measurement procedure. That makes it valuable for inference engineers. It does not make a 372 GB checkpoint convenient, nor does our 10-second Python build prove the C runtime. Adopt it as a measured systems project with pinned weights and raw logs, not as a drop-in local chat application.

