mrkeyoor.com_
Tue 06 Oct 15:53 UTC
LLM Toolsevaluationupdated 06 Oct 2026

Rapid-MLX review

Rapid-MLX is a local inference server for running language and media models on Apple Silicon, with OpenAI-compatible and Anthropic-compatible APIs. It is aimed especially at coding agents, adding tool-call parsers, prompt caching, continuous batching, model recommendations, and one-command client setup around Apple's MLX stack.

Verdict

Our Rapid-MLX run installed 212 packages and used 6,680 MB, then passed 6,510 tests but still ended with 46 failures and 154 collection/setup errors. On an Apple Silicon Mac, it is one of the most ambitious local backends for coding agents, especially if you need both OpenAI and Anthropic request formats. Choose it for that breadth, then test your exact model, parser, and client on the Mac that will serve them; choose Ollama or a narrower MLX tool when portability or a smaller operating surface matters more.

We ran it

Lab card: what happened when we ran Rapid-MLXScreenshot of Rapid-MLX (rapidmlx.com)
Install✓ · 163s212 packages · 6680 MB
Build✓ · 12s
Tests✗ · 287s6510 passed · 46 failed · 137 skipped · 154 errors of 6710 (pytest)
Known vulns0(pip-audit)
Repo3167 files~1,307,191 lines of source · 78.5 MB · 21 CI workflows · tests dir

Answers from our run

Does Rapid-MLX build from source?

Dependencies installed in 163 seconds (212 packages), and the build succeeded in 12 seconds. We cloned commit 243a6a6 into a clean Debian container with 3 CPUs and no project-specific setup.

Do Rapid-MLX's tests pass?

Not all of them: 6510 of 6710 passed and 46 failed when we ran the project's own test command (pytest), with 154 collection errors. Some failures need services or credentials a bare container does not have.

Does Rapid-MLX have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use Rapid-MLX?

Windows, Linux, or Intel Mac users: the supported runtime is Apple Silicon, and the desktop app has no Windows or Linux build.

What are the alternatives to Rapid-MLX?

oMLX, Ollama, mlx-lm. Our Rapid-MLX run installed 212 packages and used 6,680 MB, then passed 6,510 tests but still ended with 46 failures and 154 collection/setup errors.

Setup2/56,680 MB install, Apple-only runtime, and model-specific extras
Docs5/5Detailed install, model, client, privacy, and troubleshooting guides
Community5/53,913 stars, an October 6 push, and active issue work
Maturity3/5v0.15.6 is active, but our full Debian suite did not pass

Who it’s for

Apple Silicon owners who want one local endpoint for Codex CLI, Claude Code, Aider, or their own OpenAI-compatible client.
Developers willing to choose models by unified-memory limits and keep several gigabytes free for the Python environment before weights arrive.
Teams that need tool calling, prefix-cache reuse, and concurrent local requests more than cross-platform portability.
Mac users who want a CLI, server, and optional desktop interface from the same project.

Who it’s NOT for

Windows, Linux, or Intel Mac users: the supported runtime is Apple Silicon, and the desktop app has no Windows or Linux build.
Developers expecting a small Python utility: our install pulled 212 packages and occupied 6,680 MB before model weights.
Release gates that require every supplied test to pass in a fresh Debian container: our run ended with 46 failures and 154 collection/setup errors.
Users who need the audio routes to work from the base environment we tested: several failures reported that mlx-audio was not installed, while others could not load libmlx.so.
Qwen Code users with hand-edited provider lists who cannot review setup changes: open issue 4090 says the setup command replaces existing entries without a backup.
Operators who do not want default metadata telemetry and cannot enforce the documented opt-out.

Setup reality

Our sandbox installed commit 243a6a6 in 163 seconds, adding 212 packages and using 6,680 MB. The build succeeded in 12 seconds. Tests failed in 287 seconds: pytest reported 6,510 passed, 46 failed, 137 skipped, and 154 collection/setup errors out of 6,710. Pip-audit found 0 known vulnerabilities.

The supported serving path needs an Apple Silicon Mac, Python 3.10 or newer for pip installs, and enough unified memory for the chosen model. First use downloads weights. Text is the base install; vision, audio, images, video, embeddings, and other model families use optional extras. A public endpoint needs RAPID_MLX_API_KEY.

Our Debian container was useful for packaging and broad tests, not MLX inference. The log tail repeatedly reported a missing libmlx.so; audio route failures also said mlx-audio was absent. Those messages explain what the suite encountered, but they do not measure generation speed or behavior on a supported Mac.

Two API formats make Rapid-MLX unusually agent-friendly

Rapid-MLX serves OpenAI chat completions, Responses, and Anthropic Messages from one local process. That matters for coding tools because Codex CLI and Claude Code do not speak the same wire format. The project also includes 27 tool-call parser modules, an automatic fallback, continuous batching, a radix prefix cache, and disk restoration of cached state. Its compatibility guide names 12 agent clients and 3 Python frameworks, with 5 agents treated as release-blocking integrations.

The scope now reaches far beyond text generation. Optional packages add vision, embeddings, speech, transcription, voice cloning, image generation, video generation, and typed decision models. A desktop app manages many of those jobs, while the CLI covers model download, serving, health checks, client configuration, benchmarks, and telemetry. That breadth is convenient on one capable Mac. It also explains why this repository contains about 1,307,191 source lines and why a simple-looking local server has a large dependency and testing surface.

The 6,680 MB install arrives before any model weights

Our checkout was 78.5 MB across 3,167 files, but installing commit 243a6a6 pulled 212 packages and expanded the environment to 6,680 MB. The project README describes the base text install as roughly 460 MB for ordinary users, so our source-development environment should not be read as the size of a Homebrew bottle. It is still the relevant cost for contributors who plan to build and run the repository's checks in an isolated Python environment.

Model storage sits on top of that. The quick start downloads about 3 GB for its default first-chat model, and the hardware table ranges from an 8 GB memory tier to models intended for 32 GB or more. Image and video aliases can add much larger downloads. Rapid-MLX does give buyers concrete memory recommendations and a doctor command, but no catalog can predict what else is consuming unified memory on your Mac. Capacity planning remains part of the job.

What happened when we ran it

Our Debian sandbox installed commit 243a6a6 in 163 seconds and built it successfully in 12 seconds. It had 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit found 0 known vulnerabilities. The repository included 21 CI workflow files and a tests directory, but no Dockerfile. These results cover package installation and the supplied build path, not model loading or token generation on Apple hardware.

The test step failed after 287 seconds. Pytest reported 6,510 passed, 46 failed, 137 skipped, and 154 collection/setup errors out of 6,710. Several failures in the log tail raised ImportError because libmlx.so could not be opened. Audio route cases returned errors saying mlx-audio was not installed, and cache route tests hit the same missing MLX library. The log does not establish how those cases behave on a supported Apple Silicon machine.

Debian exposed the portability boundary, not Mac inference speed

Rapid-MLX explicitly targets M-series Macs, and its desktop app has no Windows or Linux build. Our Debian result therefore answers a narrower question: the Python project installs and builds there, but its complete suite does not pass without the MLX runtime and optional audio pieces named in the errors. We did not load weights, serve a request, or verify the project's performance comparisons. Buyers should treat every published speed figure as the project's measurement until they repeat it on their own Mac and workload.

The boundary is sharper than in a cross-platform local runner. Apple Silicon lets Rapid-MLX use unified memory and native MLX kernels, while tying deployment to one hardware family. Ollama is the easier answer for a mixed Windows, Linux, and macOS fleet. Apple's mlx-lm is the smaller conceptual base for developers who want generation and fine-tuning without Rapid-MLX's client setup, desktop, media routes, and broad model catalog.

Client setup can overwrite a Qwen Code provider list

The agent integrations save hand editing, but they deserve a preview before approval. Open issue 4090 says agents qwen-code --setup replaces the existing OpenAI provider list and writes no backup. Issue 4037 documents a parser turning an integer tool parameter into a float, which Codex CLI rejected in 3 of 3 reported runs. These are precise integration defects in the area Rapid-MLX emphasizes most. Test the selected model parser with real tool calls, not just a successful health response.

Authentication matters if the endpoint leaves localhost. The README tells tunnel users to set RAPID_MLX_API_KEY and warns never to expose an unauthenticated server. Metadata telemetry is enabled by default from version 0.15.0, with documented CLI and environment-variable opt-outs. The project says prompts, completions, paths, IP addresses, locations, and API keys are excluded. Privacy-sensitive deployments should still set the opt-out explicitly and verify the bind address before connecting an agent with repository access.

Version 0.15.6 is active, while 39 items remain open

Release v0.15.6 shipped on October 4, 2026, and the repository was pushed again on October 6. GitHub listed 39 open issues and pull requests: 30 issues and 9 pull requests. That is a busy maintenance picture, reinforced by 21 CI workflow files and the 6,510 passing tests in our run. The failed remainder prevents a clean endorsement from our Debian box, but it does not resemble an abandoned repository.

Rapid-MLX makes sense when one Apple Silicon machine will be a serious local agent server and you want both major API shapes, model guidance, and tool parsing in one package. Start with one text model and one client. Add media extras only when a real job requires them. The 6,680 MB development environment, optional runtimes, default telemetry, and client-specific bugs are the price of its wide scope, not details to discover after you have wired every editor to it.

Alternatives

ProjectWhat it isPick it when
oMLX gh↗An MLX inference server with continuous batching, SSD caching, and a macOS menu-bar app.pick this instead when a Mac-native menu-bar workflow and tiered SSD cache are higher priorities than Rapid-MLX's agent integrations.
Ollama gh↗A local model runner with its own API plus OpenAI and Anthropic compatibility.pick this instead when you need Windows or Linux support, GGUF models, or a larger general-user ecosystem.
mlx-lmApple's MLX-focused package for generating and fine-tuning language models.pick this instead when you want the upstream MLX language-model toolkit and can build the agent-facing service layer yourself.

What people are saying

  1. [github-trending] raullenchai/Rapid-MLX

Sources

  1. Rapid-MLX README
  2. Rapid-MLX v0.15.6 release
  3. Commit 243a6a6
  4. Issue 4090: Qwen Code setup overwrites providers
  5. Issue 4037: Codex rejects converted tool argument

More llm tools reviews

llm-d-router · simple-jev · kev · openjev-sglang · OptMem · SemIf-OpenJev · the whole board →