Two API formats make Rapid-MLX unusually agent-friendly
Rapid-MLX serves OpenAI chat completions, Responses, and Anthropic Messages from one local process. That matters for coding tools because Codex CLI and Claude Code do not speak the same wire format. The project also includes 27 tool-call parser modules, an automatic fallback, continuous batching, a radix prefix cache, and disk restoration of cached state. Its compatibility guide names 12 agent clients and 3 Python frameworks, with 5 agents treated as release-blocking integrations.
The scope now reaches far beyond text generation. Optional packages add vision, embeddings, speech, transcription, voice cloning, image generation, video generation, and typed decision models. A desktop app manages many of those jobs, while the CLI covers model download, serving, health checks, client configuration, benchmarks, and telemetry. That breadth is convenient on one capable Mac. It also explains why this repository contains about 1,307,191 source lines and why a simple-looking local server has a large dependency and testing surface.
The 6,680 MB install arrives before any model weights
Our checkout was 78.5 MB across 3,167 files, but installing commit 243a6a6 pulled 212 packages and expanded the environment to 6,680 MB. The project README describes the base text install as roughly 460 MB for ordinary users, so our source-development environment should not be read as the size of a Homebrew bottle. It is still the relevant cost for contributors who plan to build and run the repository's checks in an isolated Python environment.
Model storage sits on top of that. The quick start downloads about 3 GB for its default first-chat model, and the hardware table ranges from an 8 GB memory tier to models intended for 32 GB or more. Image and video aliases can add much larger downloads. Rapid-MLX does give buyers concrete memory recommendations and a doctor command, but no catalog can predict what else is consuming unified memory on your Mac. Capacity planning remains part of the job.
What happened when we ran it
Our Debian sandbox installed commit 243a6a6 in 163 seconds and built it successfully in 12 seconds. It had 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit found 0 known vulnerabilities. The repository included 21 CI workflow files and a tests directory, but no Dockerfile. These results cover package installation and the supplied build path, not model loading or token generation on Apple hardware.
The test step failed after 287 seconds. Pytest reported 6,510 passed, 46 failed, 137 skipped, and 154 collection/setup errors out of 6,710. Several failures in the log tail raised ImportError because libmlx.so could not be opened. Audio route cases returned errors saying mlx-audio was not installed, and cache route tests hit the same missing MLX library. The log does not establish how those cases behave on a supported Apple Silicon machine.
Debian exposed the portability boundary, not Mac inference speed
Rapid-MLX explicitly targets M-series Macs, and its desktop app has no Windows or Linux build. Our Debian result therefore answers a narrower question: the Python project installs and builds there, but its complete suite does not pass without the MLX runtime and optional audio pieces named in the errors. We did not load weights, serve a request, or verify the project's performance comparisons. Buyers should treat every published speed figure as the project's measurement until they repeat it on their own Mac and workload.
The boundary is sharper than in a cross-platform local runner. Apple Silicon lets Rapid-MLX use unified memory and native MLX kernels, while tying deployment to one hardware family. Ollama is the easier answer for a mixed Windows, Linux, and macOS fleet. Apple's mlx-lm is the smaller conceptual base for developers who want generation and fine-tuning without Rapid-MLX's client setup, desktop, media routes, and broad model catalog.
Client setup can overwrite a Qwen Code provider list
The agent integrations save hand editing, but they deserve a preview before approval. Open issue 4090 says agents qwen-code --setup replaces the existing OpenAI provider list and writes no backup. Issue 4037 documents a parser turning an integer tool parameter into a float, which Codex CLI rejected in 3 of 3 reported runs. These are precise integration defects in the area Rapid-MLX emphasizes most. Test the selected model parser with real tool calls, not just a successful health response.
Authentication matters if the endpoint leaves localhost. The README tells tunnel users to set RAPID_MLX_API_KEY and warns never to expose an unauthenticated server. Metadata telemetry is enabled by default from version 0.15.0, with documented CLI and environment-variable opt-outs. The project says prompts, completions, paths, IP addresses, locations, and API keys are excluded. Privacy-sensitive deployments should still set the opt-out explicitly and verify the bind address before connecting an agent with repository access.
Version 0.15.6 is active, while 39 items remain open
Release v0.15.6 shipped on October 4, 2026, and the repository was pushed again on October 6. GitHub listed 39 open issues and pull requests: 30 issues and 9 pull requests. That is a busy maintenance picture, reinforced by 21 CI workflow files and the 6,510 passing tests in our run. The failed remainder prevents a clean endorsement from our Debian box, but it does not resemble an abandoned repository.
Rapid-MLX makes sense when one Apple Silicon machine will be a serious local agent server and you want both major API shapes, model guidance, and tool parsing in one package. Start with one text model and one client. Add media extras only when a real job requires them. The 6,680 MB development environment, optional runtimes, default telemetry, and client-specific bugs are the price of its wide scope, not details to discover after you have wired every editor to it.

