mrkeyoor.com_
Mon 05 Oct 07:16 UTC
AI Toolsevaluationupdated 05 Oct 2026

OrcaBonsai-27B-Uncensored review

OrcaBonsai is a runtime add-on that reduces refusal behavior in the separate Ternary Bonsai 2 27B model without rewriting its compressed weights. It is research tooling for Apple Silicon and a particular model pack, not a new model checkpoint or a general chat application.

Verdict

Our OrcaBonsai install pulled 93 packages and occupied 748 MB, but there was no standard test target, so this is a research kit rather than a release-ready model service. Use it if you already have the exact Bonsai pack, Apple Silicon, and a concrete ablation experiment. Keep it out of user-facing systems unless a separate safety layer can handle the behavior it is built to expose.

We ran it

Lab card: what happened when we ran OrcaBonsai-27B-UncensoredScreenshot of OrcaBonsai-27B-Uncensored (www.orcarouter.ai)
Install✓ · 35s93 packages · 748 MB
Build✓ · 5s
Testsn/ano test script
Known vulns2(pip-audit)
Repo22 files~1,191 lines of source · 10.1 MB · 0 CI workflows

Answers from our run

Does OrcaBonsai-27B-Uncensored build from source?

Dependencies installed in 35 seconds (93 packages), and the build succeeded in 5 seconds. We cloned commit 703028b into a clean Debian container with 3 CPUs and no project-specific setup.

Does OrcaBonsai-27B-Uncensored have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does OrcaBonsai-27B-Uncensored have known vulnerabilities in its dependencies?

pip-audit flagged 2 known advisories in the dependency tree at the time of our run.

Who should not use OrcaBonsai-27B-Uncensored?

NVIDIA inference teams expecting CUDA acceleration: the README says the required quantized matrix operation has no CUDA implementation.

What are the alternatives to OrcaBonsai-27B-Uncensored?

Abliterator, MLX LM, llama.cpp. Our OrcaBonsai install pulled 93 packages and occupied 748 MB, but there was no standard test target, so this is a research kit rather than a release-ready model service.

Setup2/5Install passed, but the separate pack and narrow runtimes add friction
Docs5/5The README states runtime, format, memory, and evaluation limits
Community2/5Created in September 2026 with no releases and one open PR
Maturity2/5No CI or standard test target, plus several unverified paths

Who it’s for

MLX researchers studying refusal-direction ablation on quantized language models.
Apple Silicon developers who already have the Ternary Bonsai 2 27B pack and want reversible alpha or layer experiments.
llama.cpp experimenters willing to use PrismML's fork and verify that the supplied LoRA is active.
Safety researchers who will keep access controls outside the model and inspect outputs themselves.

Who it’s NOT for

NVIDIA inference teams expecting CUDA acceleration: the README says the required quantized matrix operation has no CUDA implementation.
Ollama users: the project says Ollama has no working path for these private quantization types and adapters.
Anyone seeking a self-contained model download: you must obtain the original Bonsai pack separately and use its bundled runtime.
Product teams relying on model refusals as their safety boundary: this code is designed to remove a learned refusal direction.
Release pipelines that require standard automated checks: our harness found no test target, the repository has no CI workflows, and pip-audit reported 2 known vulnerabilities.

Setup reality

Our sandbox installed commit 703028b in 35 seconds, pulling 93 Python packages and using 748 MB on disk. The build succeeded in 5 seconds. No standard test script or target was present, so our test step was skipped; pip-audit reported 2 known vulnerabilities.

Running the intervention needs the separate prism-ml Ternary Bonsai 2 27B pack. The main route uses MLX on Apple Silicon, the pack's bundled runtime, and a snapshot path containing config.json plus runtime files. No hosted API credential is required for local inference, though downloading from Hugging Face needs network access.

The 22-file checkout is only 10.1 MB with about 1,191 source lines because it does not contain the base weights. Linux CPU inference is described as taking minutes per forward pass, CUDA lacks the required operation, and stock llama.cpp plus Ollama cannot run the compressed packs correctly.

The code changes refusal behavior without changing the 27B pack

OrcaBonsai inserts one mathematical correction into the inference path of Ternary Bonsai 2 27B. At each residual writer, it removes the component that points along a supplied refusal direction. The original compressed weights stay bit-identical, and an alpha value controls the strength at runtime. Set alpha to 0 for the original behavior or 1 for the full projection. That reversibility is the project's best idea.

This repository is not a model release. You must download prism-ml/Ternary-Bonsai-2-27B-mlx-2bit separately, then load it through the runtime bundled with that pack. The code wraps 129 write sites across attention, linear-attention, MLP, and embedding modules. Its self-check is meant to confirm that every expected site is present and that the chosen direction has been removed from the residual stream.

What happened when we ran it

Our sandbox installed commit 703028b in 35 seconds. Pip brought in 93 packages, and the environment occupied 748 MB on disk. The build step succeeded in 5 seconds inside an unprivileged Debian container with 3 CPUs, 8 GB of RAM, and no secrets. Pip-audit reported 2 known vulnerabilities. The supplied measurement does not include their packages or severity, so both need inspection before deployment.

There was no standard test script or target for our harness to run, so tests were skipped. The repository does contain scripts/test_ablation.py, which the README describes as a model-free numerical check, and a separate self-check that needs the model pack. We did not run either as part of the measured test step. The checkout itself was 22 files, about 1,191 lines of source, and 10.1 MB.

That small size is easy to misread. The base model is outside the repository, and the README's phone calculation puts its source pack at 8.005 GiB. Our 748 MB figure covers installed code dependencies, not those weights. An install that finishes on an ordinary Linux box proves the Python environment can resolve; it does not prove that the target model will run at a useful speed.

Apple Silicon is the practical path, while CUDA is missing

The quick start is written for Apple Silicon and MLX. It expects a Hugging Face snapshot directory containing config.json and the pack's runtime. An ordinary MLX loader can appear to accept the files while producing wrong computation because it misses the activation transforms required by the compressed representation. That is a nasty failure mode: a program can run and still be invalid. Use the loader path the repository specifies.

Linux has a CPU Docker route under docker/Dockerfile, although the lab's root-level scan reported no Dockerfile. The README says a forward pass for this 27B model can take minutes on CPU. CUDA is not a useful escape because the required quantized matrix multiplication has no CUDA implementation. For routine interactive generation, the Linux path is for verification rather than serving.

The GGUF route is just as specific. A rank-1 LoRA adapter is included, but the private PTQ1_0 and PQ2_0 types require PrismML's llama.cpp fork. Stock llama.cpp rejects them. The README warns that a full-precision GGUF may load under stock software yet produce nonsense because upstream ignores its Hadamard metadata. Ollama is ruled out directly.

The published evaluations are useful, with stated limits

The authors compare alpha 0 and alpha 1 using the same weights and process. They report lower refusal rates across several harmful-prompt sets and less over-refusal on benign sets. They also label the opening-phrase classifier as indicative rather than publication-grade, disclose the 64-token response budget, and exclude MMLU-Pro because most responses did not reach an answer within that budget. Those caveats make the tables more useful, not definitive.

One limit matters even before the benchmark tables. The refusal direction came from the BF16 base model, while this project applies it to a model produced through quantization-aware training. The architecture and hidden basis match, but the README says transfer quality has not been fully measured. The runtime can prove it removed the supplied vector. It cannot prove that the vector captures the same behavior after training and compression.

Removing refusals transfers safety work to the application

The repository calls itself uncensored because it is designed to make the model answer prompts it previously declined. That can help researchers separate refusal behavior from model capability. In a product, it also removes one imperfect boundary without replacing it. Access control, prompt policy, output screening, abuse monitoring, and human review remain outside this package.

A tunable alpha does not make unsafe output predictable. The README says stronger values can degrade or collapse output, while the published harmful-prompt results still contain some refusals at full strength. Treat every prompt as an observation, not proof that a setting is safe or effective. The clean experiment is an isolated evaluation with a declared prompt set and external controls.

A September 22 push shows a new experiment, not a settled tool

GitHub says the repository was created September 18, 2026 and last pushed September 22. It has 567 stars, no published GitHub release, no CI workflow, and no tests directory. The single open tracker item is a pull request about replies being cut off at 256 tokens. Those dates show recent attention, but only across a few days.

Try OrcaBonsai when its narrow question is your question: can runtime projection alter refusal behavior while preserving a compressed pack? The code, direction files, adapter, and unusually candid README make that experiment accessible. If you need a dependable local assistant, portable inference, or an application safety boundary, choose a conventional runtime and a model whose supported hardware matches yours.

Alternatives

ProjectWhat it isPick it when
AbliteratorA Python library for feature ablation in models supported by TransformerLens.pick this instead when you want a broader ablation workflow rather than one Bonsai-specific runtime direction.
MLX LMApple's toolkit for running and fine-tuning language models with MLX.pick this instead when ordinary Apple Silicon inference matters more than refusal ablation.
llama.cpp gh↗A widely used C and C++ runtime for local language-model inference.pick this instead when your model uses upstream-supported formats and portable inference is the priority.

What people are saying

  1. [velocity-scout] Continuum-AI-Corp/OrcaBonsai-27B-Uncensored

Sources

  1. OrcaBonsai repository and technical README
  2. Ternary Bonsai 2 27B MLX 2-bit model card
  3. OrcaBonsai direction metadata
  4. Open pull request 1: reply length limit

More ai tools reviews

jev-experiments · NanoJev · uplifting-biomolecular-modeling · procedural-film · jev-review · Dream-RSI · the whole board →