mrkeyoor.com_
Thu 01 Oct 19:41 UTC
AI Toolsevaluationupdated 26 Aug 2026

mlx-audio review

MLX-Audio is a collection of speech and music inference code built on Apple's MLX framework. It gives Apple Silicon users one Python package for text to speech, transcription, speech processing, music generation, a local web interface, and an OpenAI-compatible API.

+17stars / 7d
Verdict

Our MLX-Audio run built in 13 seconds, but only 29 of 194 test outcomes passed before libmlx.so failures dominated the unsupported Debian environment. Use it on Apple Silicon when one package covering many local audio models saves more effort than maintaining separate ports. Do not select it for a Linux service, and qualify the exact model and API path rather than treating the catalog as one uniform product.

We ran it

Lab card: what happened when we ran mlx-audioScreenshot of mlx-audio (blaizzy.github.io/mlx-audio)
Install✓ · 51s106 packages · 476 MB
Build✓ · 9s
Tests✗ · 42s29 passed · 40 failed · 18 skipped · 157 errors of 226 (pytest)
Known vulns2(pip-audit)
Repo1039 files~227,621 lines of source · 15.9 MB · 5 CI workflows · tests dir

Answers from our run

Does mlx-audio build from source?

Dependencies installed in 51 seconds (106 packages), and the build succeeded in 9 seconds. We cloned commit 4ab7e6f into a clean Debian container with 3 CPUs and no project-specific setup.

Do mlx-audio's tests pass?

Not all of them: 29 of 226 passed and 40 failed when we ran the project's own test command (pytest), with 157 collection errors. Some failures need services or credentials a bare container does not have.

Does mlx-audio have known vulnerabilities in its dependencies?

pip-audit flagged 2 known advisories in the dependency tree at the time of our run.

Who should not use mlx-audio?

Linux or Windows deployments: the README requires an Apple Silicon Mac, and our Debian test collection stopped because libmlx.so was unavailable.

What are the alternatives to mlx-audio?

Coqui TTS, faster-whisper, AudioCraft. Our MLX-Audio run built in 13 seconds, but only 29 of 194 test outcomes passed before `libmlx.

Setup3/5Simple on the named Mac target; our Debian runtime could not load MLX
Docs4/5Broad examples, though model-specific behavior remains uneven
Community4/57,793 stars with an August push and active reports
Maturity3/5v0.5.0 is active, but the wide backend surface raises variance

Who it’s for

Apple Silicon developers who want several local audio model families behind one Python interface.
Mac-based teams comparing TTS, STT, diarization, source separation, and music models without moving to CUDA.
iOS or macOS builders willing to use the separate Swift package for on-device speech.

Who it’s NOT for

Linux or Windows deployments: the README requires an Apple Silicon Mac, and our Debian test collection stopped because libmlx.so was unavailable.
Teams that need one stable behavior across every listed model: individual backends have separate arguments, language support, dependencies, and model READMEs.
API users assuming every voice option maps cleanly to cloning: issue #557 reports that Qwen3 CustomVoice requests ignored supplied clone references and entered the preset-speaker path.
Buyers who cannot review model licenses and consent rules separately: the MIT license covers this code, not every downloaded voice, checkpoint, or cloned speaker.

Setup reality

Our sandbox installed 58 packages in 64 seconds and used 298 MB. The build passed in 13 seconds. Tests failed after 40 seconds: 29 passed, 39 failed, 16 skipped, and 126 collection or setup errors out of 194; the log tail says libmlx.so could not be opened.

Normal use needs Python 3.10 or newer, an Apple Silicon Mac, MLX, and model downloads from Hugging Face. Extra server dependencies are needed for the web API, while ffmpeg is required for encoded formats other than WAV.

The package spans many independent model ports, so memory, language, voices, reference-audio rules, and optional dependencies vary by backend. Our Debian container could inspect packaging and build the project, but it could not validate the Apple runtime that the README requires.

One package covers speech, transcription, and music on Apple chips

MLX-Audio packages many audio-model ports around Apple's MLX array framework. Its catalog includes text to speech, speech recognition, diarization, source separation, enhancement, audio understanding, and music generation. The common value is hardware placement: Mac developers can try these jobs without a CUDA workstation or cloud endpoint. The cost is that a catalog is not a single model. Each entry has its own weights, languages, controls, memory use, and output behavior.

The basic Python shape is approachable. Load a named model from Hugging Face, call its generation method, and consume audio arrays or transcription objects. Command-line wrappers can write or play files, stream speech, and join generated segments. A local FastAPI server exposes speech and transcription routes shaped like OpenAI's audio API. That can ease client integration, but compatibility at the URL level does not mean every model accepts the same voice, cloning, or language arguments.

Apple Silicon is a hard requirement, not a preference

The README lists Python 3.10 or newer, an Apple Silicon Mac from M1 through M4, and the MLX framework as requirements. That excludes ordinary Linux servers, Intel Macs, and Windows machines from the supported runtime. The repository does mention Debian installation for ffmpeg, but ffmpeg only handles encoded audio formats. It does not turn the MLX inference code into a Linux backend. WAV output avoids ffmpeg; it does not avoid MLX.

This focus is reasonable. A project tuned for unified memory and Apple's GPU can expose models that otherwise require separate conversion work. It also changes procurement: a service must run on Mac hardware or a Mac hosting provider, and container habits from CUDA Linux do not transfer directly. Before choosing MLX-Audio for an API, decide how the Mac host will be patched, monitored, restarted, and scaled. The package does not supply that operating layer.

What happened when we ran it

Our sandbox cloned commit ebf44df and installed 58 Python packages in 64 seconds. Dependencies used 298 MB. The source checkout contained 944 files, about 208,172 lines, and occupied 15 MB. The build succeeded in 13 seconds, and pip-audit reported 0 known vulnerabilities. The repository had 4 CI workflows and a tests directory, though no Dockerfile.

The test step failed with exit code 1 after 40 seconds. Pytest reported 29 passed, 39 failed, 16 skipped, and 126 collection or setup errors out of 194. The tail shows imports entering mlx_audio.stt.utils, then failing at import mlx.core as mx because libmlx.so could not be opened. This is consistent with running an Apple-focused library in our fresh Debian container, but the log alone does not identify every failed test.

Our results therefore say two things. Packaging and build mechanics worked on Linux, while most runtime tests did not. They do not measure speech quality, transcription accuracy, generation speed, memory use on a Mac, or whether a particular model loads. A useful evaluation needs the target Apple chip, pinned weights, sample audio, and cases drawn from the languages and file formats the product will receive.

Model breadth shifts testing to the adopter

The README lists many TTS families, from compact Kokoro and KittenTTS variants to multi-billion-parameter models with cloning or voice design. STT choices include Whisper ports, Qwen, Parakeet, Voxtral, and diarization-capable models. Music and source-separation paths use different APIs again. This is excellent for comparison work. It is harder for a stable product because fixes in one backend may not touch another, and shared CLI flags can hide model-specific meaning.

Pick one model before judging the library. Create fixtures for short text, long paragraphs, numbers, names, silence, noisy input, unsupported languages, corrupt files, and streaming cancellation. For cloning, include consent checks and verify that reference text matches reference audio where the model expects both. Save the exact Hugging Face revision. A package upgrade and a weight update are two separate changes, and either can alter results.

The OpenAI-style API still has model-specific branches

The server exposes /v1/audio/speech and /v1/audio/transcriptions, which makes standard clients a plausible starting point. Issue #557 reports a sharper edge: with a Qwen3 CustomVoice model, supplied ref_audio and ref_text were bypassed because the code entered the preset-speaker branch first. The request could therefore look like a cloning call while behaving as speaker selection. That is an open issue report, not evidence about every TTS backend.

It is enough to justify route-level tests. Send the same request through Python and HTTP, confirm the chosen speaker or reference actually changes the waveform, and make invalid combinations fail clearly. Audio endpoints often stream a 200 response before later work errors, so clients must also handle mid-stream failure. Release v0.5.0 specifically changed NDJSON streaming errors to report them in-band rather than returning an empty success response.

Active releases are adding models faster than uniformity can follow

GitHub showed 7,793 stars, 98 open issues and pull requests, and a last push on August 26, 2026. Version 0.5.0 was released August 17 with new TTS and music support, vendored MLX-LM components, hotword work, and the streaming-error change. This is an active project with a rapidly expanding catalog. Version numbers and model lists will move, so production users should pin both package and checkpoint revisions.

MLX-Audio is one of the more practical ways to turn an Apple Silicon machine into an audio-model workbench. Our 13-second build supports trying the code, while the failed Linux suite confirms that meaningful qualification belongs on a Mac. Choose it for a specific model and workflow, test that path through the interface you will ship, and resist promising the entire catalog as one interchangeable audio service.

Alternatives

ProjectWhat it isPick it when
Coqui TTSA broad speech toolkit with pretrained models, training code, and voice conversion.pick this instead when cross-platform PyTorch support and training workflows matter more than Apple MLX optimization, and an archived repository is acceptable.
faster-whisperA focused Whisper transcription implementation using CTranslate2.pick this instead when transcription is the only job and you want a narrower server or library.
AudioCraftMeta's research code for audio and music generation models such as MusicGen.pick this instead when music generation research matters more than one Mac-native speech toolbox.

What people are saying

  1. [github-trending] Blaizzy/mlx-audio

Sources

  1. MLX-Audio README
  2. MLX-Audio repository
  3. MLX-Audio v0.5.0 release
  4. Issue #557: Qwen3 cloning path
  5. Issue #420: uv and NumPy installation report

More ai tools reviews

UniMate · unigit-ecosystem · handraw-style · souchastnik · dlssg_for_sm86 · glm-flash-offline-client · the whole board →