mrkeyoor.com_
Tue 29 Sept 15:34 UTC
AI Toolsevaluationupdated 29 Sept 2026

Bonsai-demo review

Bonsai Demo packages the scripts, web interface, and custom runtimes needed to run PrismML's ternary Bonsai 2 27B model on a local computer. It gives Mac, Linux, and Windows users one place to download the model, start an OpenAI-style server, send images, call tools, and connect MCP servers.

Verdict

Our Bonsai Demo run passed 38 of 38 tests in 13 seconds, but it did not download or evaluate the 27B model, so the green suite proves the tooling rather than the model's answers. Use it if Bonsai 2 is the point and you are comfortable matching a custom runtime to your hardware. Choose Ollama or upstream llama.cpp when predictable model support matters more than this ternary model.

We ran it

Lab card: what happened when we ran Bonsai-demoScreenshot of Bonsai-demo (prismml.com)
Install✓ · 33s50 packages · 150 MB
Build✓ · 7s
Tests✓ · 13s38 passed · 0 failed of 38 (pytest)
Known vulns0(pip-audit)
Repo170 files~5,504 lines of source · 33.8 MB · 4 CI workflows · tests dir

Answers from our run

Does Bonsai-demo build from source?

Dependencies installed in 33 seconds (50 packages), and the build succeeded in 7 seconds. We cloned commit 69c3a8b into a clean Debian container with 3 CPUs and no project-specific setup.

Do Bonsai-demo's tests pass?

Yes: 38 of 38 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does Bonsai-demo have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use Bonsai-demo?

Operators who require stock llama.cpp: all Bonsai 2 formats still need PrismML's fork, and the support guide warns that one upstream format can load yet produce bad output.

What are the alternatives to Bonsai-demo?

Ollama, llama.cpp, MLX LM. Our Bonsai Demo run passed 38 of 38 tests in 13 seconds, but it did not download or evaluate the 27B model, so the green suite proves the tooling rather than the model's answers.

Setup3/5Scripts help, but the model is large and runtime support is fork-specific
Docs5/5Detailed backend matrix, hardware knobs, setup guide, and FAQ
Community4/53,199 stars, a September 28 push, and active issue reports
Maturity3/5Tests pass, while runtime and client compatibility remain in motion

Who it’s for

Local-model enthusiasts with enough RAM or GPU memory to tune a 27B model.
Developers testing ternary weights, long context, vision, or tool calls on their own hardware.
Apple Silicon users willing to work with MLX and PrismML's pinned runtime.
Teams prepared to diagnose backend, model-format, and chat-template compatibility.

Who it’s NOT for

Operators who require stock llama.cpp: all Bonsai 2 formats still need PrismML's fork, and the support guide warns that one upstream format can load yet produce bad output.
People expecting a small download: the default setup fetches 7.8 GB of model weights and a vision projector, then adds several GB if Open WebUI and the code interpreter stay enabled.
Apple users relying on documented MLX thinking today: open issue 245 reports that the chat template is not applied and reasoning_content stays empty.
Claude Code or Anthropic-protocol users who cannot patch around chat-template failures: open issue 213 reports HTTP 500 responses for mid-conversation system messages.
Teams that need the hardware defaults to be dependable across GPUs: open issue 193 says the RAM-based context choice can exceed a mid-size GPU's VRAM.

Setup reality

Our sandbox installed commit 69c3a8b in 33 seconds, adding 50 packages and 150 MB. The build passed in 7 seconds. Pytest then passed all 38 tests in 13 seconds, and pip-audit reported 0 known vulnerabilities.

A full model setup is much larger than our dependency check. The README says the default script downloads 7.8 GB of Bonsai 2 weights plus a vision projector. Open WebUI and the code-interpreter environment add several GB, though both can be disabled. Public model files need no token.

Bonsai 2 needs PrismML's llama.cpp fork or its pinned MLX path. Backend support varies by model packing and hardware, and the repository has no Dockerfile. Current issues document MLX, CUDA, Vulkan, context sizing, and chat-template failures.

Bonsai 2 turns a 27B model into a 7.8 GB local download

Bonsai Demo is the delivery kit for PrismML's ternary Bonsai 2 27B model. The default setup fetches a 7.8 GB PQ2_0 package plus its vision projector, installs the matching runtime, and can add Open WebUI with a code interpreter. The resulting local server accepts text and images, exposes OpenAI-style tool calls, and lets the web interface opt into MCP servers per chat.

That is an appealing amount of model in a laptop-oriented package, but this repository is not the model evaluation. PrismML publishes quality claims and community hardware results elsewhere. Our sandbox did not score answers, time token generation, exercise vision, or run an agent loop. A buyer should separate the convenience of these scripts from the harder question of whether Bonsai 2 is good enough for the intended work.

Every Bonsai 2 format still depends on PrismML's fork

The project states that stock llama.cpp cannot correctly run Bonsai 2 today. Its PQ2_0 and PTQ1_0 packings are fork-specific. The upstream Q2_0 type can recognize and load a development model while omitting Bonsai's required transforms, which the support guide says produces gibberish. A successful load is therefore not a sufficient compatibility check.

PrismML pins its own llama.cpp binaries for CPU, Metal, CUDA, ROCm, Vulkan, and Windows variants. The backend table is unusually frank about gaps. PQ2_0, the default download, has no native Vulkan kernels, while SYCL paths still need validation. The table also says source-level support does not guarantee a particular device or configuration. You have to match the model packing, backend, driver, and effective launch flags.

What happened when we ran it

Our sandbox installed commit 69c3a8b in 33 seconds, adding 50 Python packages and using 150 MB on disk. The build completed in 7 seconds. Pytest then finished in 13 seconds with 38 passed and 0 failed. Pip-audit reported 0 known vulnerabilities in the installed Python environment.

The checkout contained 170 files, about 5,504 lines of source, and occupied 33.8 MB before installation. It had 4 CI workflow files and a tests directory, but no Dockerfile. That is a compact, tested control repository compared with the runtime assets it orchestrates. Keeping the scripts separate from the model files also makes review and updates easier.

Those 38 passing tests do not include model inference in our measured result. We did not fetch the 7.8 GB weight package, start the custom llama.cpp server, or inspect generated output. The result supports a narrow claim: the Python project installed, built, and passed its supplied suite in a fresh Debian container with 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges.

RAM-based context sizing can overflow GPU memory

The launch scripts choose a context tier from system RAM, ranging from 8K on smaller machines to 131K on machines above the documented threshold. Open issue 193 shows why that shortcut can fail. A computer may have enough system memory for a large tier while its GPU lacks space for the model, vision projector, compute buffers, and KV cache at once.

The README exposes BONSAI_CTX, layer offload, and a 4-bit KV option so an experienced user can correct the choice. The FAQ also covers startup freezes, Metal failures, CUDA detection, and source-build memory pressure. These are useful controls, but their presence tells you what kind of product this is: a hardware-aware demo where the operator is expected to read logs and tune memory.

Current MLX and Anthropic paths have open failures

Open issue 245 reports that the native MLX server can answer while failing to apply its chat template, leaving the thinking phase unopened and reasoning_content empty. Issue 246 says setup.sh stops the MLX step without full Xcode even though the reported Bonsai 2 wheel path works with command-line tools. Both were still open on September 29, 2026.

A separate issue, 213, reports that the packaged chat template rejects a system message anywhere after the first position. The reporter reproduced HTTP 500 responses with Anthropic-style clients, including Claude Code workflows that insert system reminders during a conversation. Another open report covers RAM-selected context exceeding VRAM. These are specific integration limits, not vague early-project caution.

September activity is fast, with 15 issues still open

GitHub showed 3,199 stars, 15 open issues, and 6 open pull requests on September 29, 2026. The measured commit landed on September 28, adding Apple hardware results, and the repository was pushed that day. GitHub returned no latest release for this demo, though it pins a release from the separate PrismML llama.cpp fork.

The pace suggests maintainers are working on real hardware reports rather than leaving users alone with them. It also means instructions and pinned binaries can change within days. Bonsai Demo deserves a trial if you specifically want Bonsai 2 and can validate output on the target machine. If the model itself is interchangeable, a mainstream runner avoids the fork, format matrix, and client-template work documented here.

Alternatives

ProjectWhat it isPick it when
Ollama gh↗A local model runner with a broad model catalog and a simple service interface.pick this instead when easy model switching matters more than running Bonsai 2 specifically.
llama.cpp gh↗The upstream C and C++ inference runtime behind many local-model tools.pick this instead when upstream model support and runtime control matter more than PrismML's fork-only formats.
MLX LMApple's MLX tools for running and fine-tuning language models on Apple Silicon.pick this instead when you use a supported MLX language model and do not need Bonsai's vision demo.

What people are saying

  1. [github-trending] PrismML-Eng/Bonsai-demo

Sources

  1. Bonsai Demo README
  2. Bonsai backend support matrix
  3. Bonsai troubleshooting FAQ
  4. Open MLX thinking issue
  5. Open Anthropic chat-template issue
  6. Open GPU context-sizing issue
  7. Bonsai Demo repository facts

More ai tools reviews

voltagent · InferenceX · qwen-audio-agent · wechat-intelligence-hub · dlss5-visual-enhancer · ABot-Recon · the whole board →