Laya-MLX runs typed decisions only on Apple Silicon
Laya-MLX takes state plus one or more typed questions and returns probabilities for named choices, ordered score levels, or a true proposition. It does not produce prose or JSON token by token. The port implements Laya's ModernBERT encoder, decision transformer, scoring head, and action head in Apple's MLX framework. That gives a Mac application a local decision component without PyTorch, a Transformers runtime, or a cloud request after the model files have been downloaded.
Three checkpoints are supported: a 421M-parameter English model with a 512-token context, a 322M multilingual model with 1,024 tokens, and a 421M typed-decisions model with 1,024 tokens. The project publishes converted FP16 weights on Hugging Face and can convert pinned upstream checkpoints. Training and fine-tuning remain upstream. Laya-MLX is explicit that it is an independent port, not an official Convai Innovations release and not a complete copy of the upstream API.
The 150 MB environment does not include model weights
Our checkout at commit ca5940a contained 132 files, roughly 9,224 lines of source, and occupied 24.2 MB. Installation pulled 51 packages in 17 seconds and used 150 MB on disk. Those numbers stop before a checkpoint download. The repository has 2 CI workflow files and a tests directory, but no Dockerfile. For the intended platform, the README requires Apple Silicon, Python 3.11 or newer, and macOS 14 or newer.
First load downloads the selected checkpoint; later calls can remain local. The Python API accepts text, dictionaries, or conversation lists as state. A router can keep up to 2 models loaded, choose the multilingual checkpoint from language signals, or select a task-specific checkpoint explicitly. That convenience costs memory, especially when models stay resident. Published project figures list peak allocation separately for the English and multilingual checkpoints, but those M3 Max numbers were not produced by our lab.
What happened when we ran it
Our Debian sandbox installed commit ca5940a in 17 seconds and built it successfully in 5 seconds. The container used 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit found 0 known vulnerabilities among the 51 installed packages. Our measurement method covered package installation, build, the repository test command, and audit. It did not download a checkpoint or run inference.
Test collection failed after 3 seconds with exit 4. The traceback ends in tests/conftest.py, where import mlx.core as mx raised ModuleNotFoundError: No module named 'mlx'. The log does not report a failed assertion or any completed test count. Debian is outside the README's supported Apple Silicon macOS target, but the defensible lab finding is simply that a successful install and build did not make the suite runnable in this fresh container.
The M3 Max figures are project benchmarks, not our result
The README reports a 13.42 ms median for one short English question and 7.39 ms for the multilingual checkpoint on an M3 Max with 40 GPU cores and 128 GiB of memory. It says timing includes prompt preparation, tokenization, synchronized inference, calibration, and result formatting, while excluding model load. Batch throughput uses a different 50-question workload. Those boundaries are unusually clear, but the numbers should not be transferred to another Mac or a longer prompt without measurement.
Validation also receives careful treatment. The project says all 3 checkpoints matched upstream selected answers on 63 validation questions in FP16 and FP32, and it publishes raw benchmark artifacts. That is port fidelity on named fixtures, not general accuracy. The runtime even clamps one raw calibration bucket because its stored temperature could make a coin flip appear nearly certain. This is the right kind of warning: a calibrated probability still needs task-specific evaluation before it controls an action.
Shared option tokens limit very large label sets
Choice labels share one token budget. With hundreds of labels, each may receive only a few tokens, causing distinct options to collapse into similar spans. The optional shortlisting helper embeds the state and labels, keeps a top subset, and sends that subset through one decision call. It can make a large taxonomy workable, but the resulting probabilities cover only retained labels. The README says a dedicated bi-encoder usually produces a better shortlist than the decision checkpoint's own encoder.
Release v0.3.0 shipped on October 2, 2026, the same date as the latest repository push. GitHub showed 6,798 stars and 0 open issues or pull requests. The release adds selected upstream behavior while declining full parity, and the README names omitted batch, long-document, hooks, and server APIs. Laya-MLX is a focused Mac runtime with unusually inspectable evidence. Its platform boundary is firm: test it on the exact Apple hardware that will run it, or pick a different implementation.

