The code changes refusal behavior without changing the 27B pack
OrcaBonsai inserts one mathematical correction into the inference path of Ternary Bonsai 2 27B. At each residual writer, it removes the component that points along a supplied refusal direction. The original compressed weights stay bit-identical, and an alpha value controls the strength at runtime. Set alpha to 0 for the original behavior or 1 for the full projection. That reversibility is the project's best idea.
This repository is not a model release. You must download prism-ml/Ternary-Bonsai-2-27B-mlx-2bit separately, then load it through the runtime bundled with that pack. The code wraps 129 write sites across attention, linear-attention, MLP, and embedding modules. Its self-check is meant to confirm that every expected site is present and that the chosen direction has been removed from the residual stream.
What happened when we ran it
Our sandbox installed commit 703028b in 35 seconds. Pip brought in 93 packages, and the environment occupied 748 MB on disk. The build step succeeded in 5 seconds inside an unprivileged Debian container with 3 CPUs, 8 GB of RAM, and no secrets. Pip-audit reported 2 known vulnerabilities. The supplied measurement does not include their packages or severity, so both need inspection before deployment.
There was no standard test script or target for our harness to run, so tests were skipped. The repository does contain scripts/test_ablation.py, which the README describes as a model-free numerical check, and a separate self-check that needs the model pack. We did not run either as part of the measured test step. The checkout itself was 22 files, about 1,191 lines of source, and 10.1 MB.
That small size is easy to misread. The base model is outside the repository, and the README's phone calculation puts its source pack at 8.005 GiB. Our 748 MB figure covers installed code dependencies, not those weights. An install that finishes on an ordinary Linux box proves the Python environment can resolve; it does not prove that the target model will run at a useful speed.
Apple Silicon is the practical path, while CUDA is missing
The quick start is written for Apple Silicon and MLX. It expects a Hugging Face snapshot directory containing config.json and the pack's runtime. An ordinary MLX loader can appear to accept the files while producing wrong computation because it misses the activation transforms required by the compressed representation. That is a nasty failure mode: a program can run and still be invalid. Use the loader path the repository specifies.
Linux has a CPU Docker route under docker/Dockerfile, although the lab's root-level scan reported no Dockerfile. The README says a forward pass for this 27B model can take minutes on CPU. CUDA is not a useful escape because the required quantized matrix multiplication has no CUDA implementation. For routine interactive generation, the Linux path is for verification rather than serving.
The GGUF route is just as specific. A rank-1 LoRA adapter is included, but the private PTQ1_0 and PQ2_0 types require PrismML's llama.cpp fork. Stock llama.cpp rejects them. The README warns that a full-precision GGUF may load under stock software yet produce nonsense because upstream ignores its Hadamard metadata. Ollama is ruled out directly.
The published evaluations are useful, with stated limits
The authors compare alpha 0 and alpha 1 using the same weights and process. They report lower refusal rates across several harmful-prompt sets and less over-refusal on benign sets. They also label the opening-phrase classifier as indicative rather than publication-grade, disclose the 64-token response budget, and exclude MMLU-Pro because most responses did not reach an answer within that budget. Those caveats make the tables more useful, not definitive.
One limit matters even before the benchmark tables. The refusal direction came from the BF16 base model, while this project applies it to a model produced through quantization-aware training. The architecture and hidden basis match, but the README says transfer quality has not been fully measured. The runtime can prove it removed the supplied vector. It cannot prove that the vector captures the same behavior after training and compression.
Removing refusals transfers safety work to the application
The repository calls itself uncensored because it is designed to make the model answer prompts it previously declined. That can help researchers separate refusal behavior from model capability. In a product, it also removes one imperfect boundary without replacing it. Access control, prompt policy, output screening, abuse monitoring, and human review remain outside this package.
A tunable alpha does not make unsafe output predictable. The README says stronger values can degrade or collapse output, while the published harmful-prompt results still contain some refusals at full strength. Treat every prompt as an observation, not proof that a setting is safe or effective. The clean experiment is an isolated evaluation with a declared prompt set and external controls.
A September 22 push shows a new experiment, not a settled tool
GitHub says the repository was created September 18, 2026 and last pushed September 22. It has 567 stars, no published GitHub release, no CI workflow, and no tests directory. The single open tracker item is a pull request about replies being cut off at 256 tokens. Those dates show recent attention, but only across a few days.
Try OrcaBonsai when its narrow question is your question: can runtime projection alter refusal behavior while preserving a compressed pack? The code, direction files, adapter, and unusually candid README make that experiment accessible. If you need a dependable local assistant, portable inference, or an application safety boundary, choose a conventional runtime and a model whose supported hardware matches yours.

