One package covers speech, transcription, and music on Apple chips
MLX-Audio packages many audio-model ports around Apple's MLX array framework. Its catalog includes text to speech, speech recognition, diarization, source separation, enhancement, audio understanding, and music generation. The common value is hardware placement: Mac developers can try these jobs without a CUDA workstation or cloud endpoint. The cost is that a catalog is not a single model. Each entry has its own weights, languages, controls, memory use, and output behavior.
The basic Python shape is approachable. Load a named model from Hugging Face, call its generation method, and consume audio arrays or transcription objects. Command-line wrappers can write or play files, stream speech, and join generated segments. A local FastAPI server exposes speech and transcription routes shaped like OpenAI's audio API. That can ease client integration, but compatibility at the URL level does not mean every model accepts the same voice, cloning, or language arguments.
Apple Silicon is a hard requirement, not a preference
The README lists Python 3.10 or newer, an Apple Silicon Mac from M1 through M4, and the MLX framework as requirements. That excludes ordinary Linux servers, Intel Macs, and Windows machines from the supported runtime. The repository does mention Debian installation for ffmpeg, but ffmpeg only handles encoded audio formats. It does not turn the MLX inference code into a Linux backend. WAV output avoids ffmpeg; it does not avoid MLX.
This focus is reasonable. A project tuned for unified memory and Apple's GPU can expose models that otherwise require separate conversion work. It also changes procurement: a service must run on Mac hardware or a Mac hosting provider, and container habits from CUDA Linux do not transfer directly. Before choosing MLX-Audio for an API, decide how the Mac host will be patched, monitored, restarted, and scaled. The package does not supply that operating layer.
What happened when we ran it
Our sandbox cloned commit ebf44df and installed 58 Python packages in 64 seconds. Dependencies used 298 MB. The source checkout contained 944 files, about 208,172 lines, and occupied 15 MB. The build succeeded in 13 seconds, and pip-audit reported 0 known vulnerabilities. The repository had 4 CI workflows and a tests directory, though no Dockerfile.
The test step failed with exit code 1 after 40 seconds. Pytest reported 29 passed, 39 failed, 16 skipped, and 126 collection or setup errors out of 194. The tail shows imports entering mlx_audio.stt.utils, then failing at import mlx.core as mx because libmlx.so could not be opened. This is consistent with running an Apple-focused library in our fresh Debian container, but the log alone does not identify every failed test.
Our results therefore say two things. Packaging and build mechanics worked on Linux, while most runtime tests did not. They do not measure speech quality, transcription accuracy, generation speed, memory use on a Mac, or whether a particular model loads. A useful evaluation needs the target Apple chip, pinned weights, sample audio, and cases drawn from the languages and file formats the product will receive.
Model breadth shifts testing to the adopter
The README lists many TTS families, from compact Kokoro and KittenTTS variants to multi-billion-parameter models with cloning or voice design. STT choices include Whisper ports, Qwen, Parakeet, Voxtral, and diarization-capable models. Music and source-separation paths use different APIs again. This is excellent for comparison work. It is harder for a stable product because fixes in one backend may not touch another, and shared CLI flags can hide model-specific meaning.
Pick one model before judging the library. Create fixtures for short text, long paragraphs, numbers, names, silence, noisy input, unsupported languages, corrupt files, and streaming cancellation. For cloning, include consent checks and verify that reference text matches reference audio where the model expects both. Save the exact Hugging Face revision. A package upgrade and a weight update are two separate changes, and either can alter results.
The OpenAI-style API still has model-specific branches
The server exposes /v1/audio/speech and /v1/audio/transcriptions, which makes standard clients a plausible starting point. Issue #557 reports a sharper edge: with a Qwen3 CustomVoice model, supplied ref_audio and ref_text were bypassed because the code entered the preset-speaker branch first. The request could therefore look like a cloning call while behaving as speaker selection. That is an open issue report, not evidence about every TTS backend.
It is enough to justify route-level tests. Send the same request through Python and HTTP, confirm the chosen speaker or reference actually changes the waveform, and make invalid combinations fail clearly. Audio endpoints often stream a 200 response before later work errors, so clients must also handle mid-stream failure. Release v0.5.0 specifically changed NDJSON streaming errors to report them in-band rather than returning an empty success response.
Active releases are adding models faster than uniformity can follow
GitHub showed 7,793 stars, 98 open issues and pull requests, and a last push on August 26, 2026. Version 0.5.0 was released August 17 with new TTS and music support, vendored MLX-LM components, hotword work, and the streaming-error change. This is an active project with a rapidly expanding catalog. Version numbers and model lists will move, so production users should pin both package and checkpoint revisions.
MLX-Audio is one of the more practical ways to turn an Apple Silicon machine into an audio-model workbench. Our 13-second build supports trying the code, while the failed Linux suite confirms that meaningful qualification belongs on a Mac. Choose it for a specific model and workflow, test that path through the interface you will ship, and resist promising the entire catalog as one interchangeable audio service.

