Bonsai 2 turns a 27B model into a 7.8 GB local download
Bonsai Demo is the delivery kit for PrismML's ternary Bonsai 2 27B model. The default setup fetches a 7.8 GB PQ2_0 package plus its vision projector, installs the matching runtime, and can add Open WebUI with a code interpreter. The resulting local server accepts text and images, exposes OpenAI-style tool calls, and lets the web interface opt into MCP servers per chat.
That is an appealing amount of model in a laptop-oriented package, but this repository is not the model evaluation. PrismML publishes quality claims and community hardware results elsewhere. Our sandbox did not score answers, time token generation, exercise vision, or run an agent loop. A buyer should separate the convenience of these scripts from the harder question of whether Bonsai 2 is good enough for the intended work.
Every Bonsai 2 format still depends on PrismML's fork
The project states that stock llama.cpp cannot correctly run Bonsai 2 today. Its PQ2_0 and PTQ1_0 packings are fork-specific. The upstream Q2_0 type can recognize and load a development model while omitting Bonsai's required transforms, which the support guide says produces gibberish. A successful load is therefore not a sufficient compatibility check.
PrismML pins its own llama.cpp binaries for CPU, Metal, CUDA, ROCm, Vulkan, and Windows variants. The backend table is unusually frank about gaps. PQ2_0, the default download, has no native Vulkan kernels, while SYCL paths still need validation. The table also says source-level support does not guarantee a particular device or configuration. You have to match the model packing, backend, driver, and effective launch flags.
What happened when we ran it
Our sandbox installed commit 69c3a8b in 33 seconds, adding 50 Python packages and using 150 MB on disk. The build completed in 7 seconds. Pytest then finished in 13 seconds with 38 passed and 0 failed. Pip-audit reported 0 known vulnerabilities in the installed Python environment.
The checkout contained 170 files, about 5,504 lines of source, and occupied 33.8 MB before installation. It had 4 CI workflow files and a tests directory, but no Dockerfile. That is a compact, tested control repository compared with the runtime assets it orchestrates. Keeping the scripts separate from the model files also makes review and updates easier.
Those 38 passing tests do not include model inference in our measured result. We did not fetch the 7.8 GB weight package, start the custom llama.cpp server, or inspect generated output. The result supports a narrow claim: the Python project installed, built, and passed its supplied suite in a fresh Debian container with 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges.
RAM-based context sizing can overflow GPU memory
The launch scripts choose a context tier from system RAM, ranging from 8K on smaller machines to 131K on machines above the documented threshold. Open issue 193 shows why that shortcut can fail. A computer may have enough system memory for a large tier while its GPU lacks space for the model, vision projector, compute buffers, and KV cache at once.
The README exposes BONSAI_CTX, layer offload, and a 4-bit KV option so an experienced user can correct the choice. The FAQ also covers startup freezes, Metal failures, CUDA detection, and source-build memory pressure. These are useful controls, but their presence tells you what kind of product this is: a hardware-aware demo where the operator is expected to read logs and tune memory.
Current MLX and Anthropic paths have open failures
Open issue 245 reports that the native MLX server can answer while failing to apply its chat template, leaving the thinking phase unopened and reasoning_content empty. Issue 246 says setup.sh stops the MLX step without full Xcode even though the reported Bonsai 2 wheel path works with command-line tools. Both were still open on September 29, 2026.
A separate issue, 213, reports that the packaged chat template rejects a system message anywhere after the first position. The reporter reproduced HTTP 500 responses with Anthropic-style clients, including Claude Code workflows that insert system reminders during a conversation. Another open report covers RAM-selected context exceeding VRAM. These are specific integration limits, not vague early-project caution.
September activity is fast, with 15 issues still open
GitHub showed 3,199 stars, 15 open issues, and 6 open pull requests on September 29, 2026. The measured commit landed on September 28, adding Apple hardware results, and the repository was pushed that day. GitHub returned no latest release for this demo, though it pins a release from the separate PrismML llama.cpp fork.
The pace suggests maintainers are working on real hardware reports rather than leaving users alone with them. It also means instructions and pinned binaries can change within days. Bonsai Demo deserves a trial if you specifically want Bonsai 2 and can validate output on the target machine. If the model itself is interchangeable, a mainstream runner avoids the fork, format matrix, and client-template work documented here.

