mrkeyoor.com_
Tue 06 Oct 07:20 UTC
Self-Hostedevaluationupdated 06 Oct 2026

splash review

Splash is a local inference server for selected large language models on recent Apple silicon Macs. It exposes OpenAI- and Anthropic-compatible APIs, a browser chat page, vision and tool calling, plus launchers that point installed coding agents at the local model.

Verdict

Our Splash run built in 6 seconds, but its 293-test command ended with 26 failures and 47 collection/setup errors in a fresh Debian container, so the repository is not a platform-neutral development experience. Try the Homebrew release if you own an M3-or-newer Mac with at least 36 GB for the main 4-bit examples and your model is explicitly supported. Choose a broader engine if hardware reach or model choice matters more than specialized Metal work.

We ran it

Lab card: what happened when we ran splashScreenshot of splash (inco.ai/blog/splash)
Install✓ · 26s49 packages · 81 MB
Build✓ · 6s
Tests✗ · 157s220 passed · 26 failed · 20 skipped · 47 errors of 293 (pytest)
Known vulns0(pip-audit)
Repo505 files~127,966 lines of source · 8.4 MB · 2 CI workflows · tests dir

Answers from our run

Does splash build from source?

Dependencies installed in 26 seconds (49 packages), and the build succeeded in 6 seconds. We cloned commit 4c11cbf into a clean Debian container with 3 CPUs and no project-specific setup.

Do splash's tests pass?

Not all of them: 220 of 293 passed and 26 failed when we ran the project's own test command (pytest), with 47 collection errors. Some failures need services or credentials a bare container does not have.

Does splash have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use splash?

Intel Mac, M1, M2, Linux, or Windows users: the documented runtime requires Apple M3 or newer and macOS 26.4 or later.

What are the alternatives to splash?

llama.cpp, MLX LM, Ollama. Our Splash run built in 6 seconds, but its 293-test command ended with 26 failures and 47 collection/setup errors in a fresh Debian container, so the repository is not a platform-neutral development experience.

Setup2/5Homebrew is simple, but hardware limits and the Debian suite failed
Docs5/5Hardware, models, memory, API behavior, and source work are explicit
Community4/51,221 stars with active October issue and pull-request work
Maturity3/5v1.2.1 ships quickly, while compatibility remains deliberately narrow

Who it’s for

Mac developers with M3-or-newer hardware and enough unified memory for a supported model.
Coding-agent users who want prompts and inference to stay on one machine.
Teams needing OpenAI or Anthropic API shapes for a local Apple silicon service.
Performance engineers willing to work inside a narrow, model-specific Metal runtime.

Who it’s NOT for

Intel Mac, M1, M2, Linux, or Windows users: the documented runtime requires Apple M3 or newer and macOS 26.4 or later.
Standard laptops with little free unified memory: the 4-bit quick-start examples need at least 36 GB, with 48 GB recommended; 24 GB Macs need smaller GGUF variants.
Users who expect broad model compatibility: support centers on two Qwen families and Ternary Bonsai 2, while open issue 297 requests upstream MLX 8-bit support.
Contributors who require a green fresh-Debian suite: our run ended with 26 failed tests and 47 collection/setup errors.
LAN operators who want secure defaults without configuration: the server starts on localhost without authentication, and remote access needs explicit host, origin, and API-key settings.

Setup reality

Our sandbox installed commit 4c11cbf from ./dev/ in 26 seconds, pulling 49 packages and using 81 MB. The build passed in 6 seconds. Pytest exited 1 after 157 seconds: 220 passed, 26 failed, 20 skipped, and 47 collection/setup errors out of 293.

End users install through Homebrew, then download a supported target model and its matching draft on first serve. Public Hugging Face repositories need no credential; private or gated models need HF_TOKEN or a Hugging Face login. Model files require substantial additional disk space.

Source builds require Apple silicon, macOS 26.4 or later, Xcode 26 or newer, Python 3.12 through 3.14, and a suitable Metal 4 compiler. Our Debian build result therefore does not validate the native Metal runtime or the Homebrew package.

Splash turns one recent Mac into a local agent server

Splash serves a supported language model on one Apple silicon machine, then lets coding agents and API clients talk to it over familiar request formats. The server handles OpenAI Chat Completions, Responses, and Completions plus Anthropic Messages. It also has streaming, tool calls, structured JSON output, base64 images and PDFs, a browser chat page, and automatic batching for concurrent requests.

The easiest route is Homebrew, followed by splash serve with a supported Hugging Face model. The first start downloads both the target weights and a matching speculative-decoding draft. Public repositories need no login. Private or gated weights require HF_TOKEN or a Hugging Face login. Installed models are reused, while the default idle policy releases weight memory after 10 minutes without a request.

Launchers for Codex, Claude, OpenCode, Hermes, and Pi point already installed agents at the local server. They do not install those clients. This is a useful distinction: Splash replaces the inference endpoint, while each agent still owns its sessions and surrounding workflow. Compatibility with two major API styles lowers application work, though it cannot make an unsupported model fit the engine's kernels.

M3 hardware and a short model list define the fit

The packaged runtime requires Apple M3 or newer and macOS 26.4 or later. The README's main 4-bit examples need at least 36 GB of unified memory, with 48 GB recommended. A 24 GB Mac can run smaller GGUF variants. Intel Macs, M1 and M2 systems, Linux servers, and Windows workstations are outside the supported path.

Model support is equally focused. Splash names Qwen3.8-27B, Qwen3.6-35B-A3B, and Ternary Bonsai 2. MLX targets must use affine 4-bit weights with groups of 64, while selected GGUF quantizations cover a wider range. Open issue 297 asks for ordinary upstream MLX 8-bit checkpoints. If your preferred fine-tune or format is not listed, assume nothing until the compatibility rules accept it.

Storage arrives after the small source checkout. The development guide says the example target and draft downloads can consume roughly 21 GB or 24 GB, depending on the family. Splash uses the Hugging Face cache instead of keeping another full copy. Operators can cap Metal memory, context, and SSD cache, but each setting trades capacity or reuse against room for other Mac applications.

What happened when we ran it

Our fresh Debian sandbox installed commit 4c11cbf from the repository's dev/ directory in 26 seconds. Pip pulled 49 packages, and the environment used 81 MB on disk. The build succeeded in 6 seconds. Pip-audit found 0 known vulnerabilities. The checkout contained 505 files, about 127,966 lines of source, and occupied 8.4 MB.

Pytest failed after 157 seconds. The supplied summary recorded 220 passed, 26 failed, 20 skipped, and 47 collection/setup errors out of 293, plus 2,830 passing subtests. Several image tests ended with ModuleNotFoundError: No module named 'PIL'. Makefile tests also logged /usr/bin/lockf: not found, and another rebuild case showed a FileNotFoundError for a path beginning under /tmp.

Those lines identify missing pieces, not their cause. We cannot tell from the log tail whether the intended developer bootstrap should have installed Pillow or whether those makefile cases are meant to run only on macOS. The result is still useful: the documented native platform is a recent Mac, and the repository-wide Python suite did not adapt cleanly to our fresh Debian environment.

The 6-second build also needs the right boundary. Splash's source guide requires Apple silicon, Xcode 26 or newer, Python 3.12 through 3.14, and a Metal 4 compiler with a specific tensor feature. Our Python-based lab step proves the measured build target passed. It does not prove that the Metal kernels compile, a model loads, or inference runs on unsupported Debian hardware.

Localhost is safe by default, while LAN use needs work

The default server listens on 127.0.0.1:8000 without authentication. That is sensible for one Mac and one user. Exposing it to a LAN requires an explicit listen address, allowed host names, allowed browser origins, and an API key. The root chat page, health endpoint, and readiness endpoint remain public according to the development guide, so a reverse proxy policy should account for them.

Memory controls deserve equal attention. Splash can cap Metal allocations and context, store cached prompt state on SSD, and preserve that cache across restarts. Version 1.2.1 made idle weight release configurable and added status fields for releases and restores. Persistent cache files and downloaded weights can be large, so laptop operators should watch both unified memory pressure and disk use during real agent sessions.

Version 1.2.1 is active, and compatibility is moving fast

The repository was pushed on October 6, 2026, one day after version 1.2.1 shipped. GitHub showed 1,221 stars, 36 open issues, and 31 open pull requests. The release fixed tool-call argument handling, requests spanning Mac sleep, idle memory timing, and reasoning settings. That is active maintenance around real serving behavior, not a dormant performance experiment.

The project publishes performance comparisons, but our sandbox did not load a model or measure tokens per second. Buyers should rerun the supplied checks on the exact Mac, model revision, quantization, context, and agent workload they plan to use. Splash is compelling for its narrow target. Llama.cpp, MLX LM, and Ollama remain safer starting points when your hardware or model catalog falls outside that target.

Alternatives

ProjectWhat it isPick it when
llama.cpp gh↗A cross-platform C/C++ inference engine with broad GGUF model support.pick this instead when model breadth and Linux, Windows, or older Mac support matter more than Splash's specialized Apple kernels.
MLX LMApple's MLX-based toolkit for running and fine-tuning language models on Apple silicon.pick this instead when you need a broader MLX workflow or fine-tuning rather than Splash's serving and coding-agent focus.
Ollama gh↗A local model manager and API with a large catalog across macOS, Linux, and Windows.pick this instead when easy model switching and cross-platform installation matter more than a model-specific Mac runtime.

What people are saying

  1. [velocity-scout] incoai/splash
  2. [hackernews] Splash-free urinals (2025)

Sources

  1. Splash README
  2. Splash 1.2.1 release
  3. Splash development guide
  4. MLX 8-bit support request

More self-hosted reviews

quivr · bindery · esp32-c3-adblock · ALVR · hysteria · skillbox · the whole board →