mrkeyoor.com_
Wed 07 Oct 06:43 UTC
AI Toolsevaluationupdated 07 Oct 2026

laya-mlx review

Laya-MLX is a native Apple Silicon runtime for open-weight Laya models that answer typed choice, score, and probability questions without generating text. It ports the encoder and decision heads to MLX for local macOS inference and includes English plus linked Chinese documentation.

Verdict

Our laya-mlx run installed 51 packages and built in 5 seconds, but its tests stopped in 3 seconds because mlx was unavailable in Debian, confirming that this is an Apple Silicon choice rather than a portable Python runtime. On a supported Mac, the local weights, explicit validation work, and typed API make it worth testing. For a server fleet or mixed developer machines, use upstream Laya or a hosted decision API.

We ran it

Lab card: what happened when we ran laya-mlxScreenshot of laya-mlx (pypi.org/project/laya-mlx)
Install✓ · 17s51 packages · 150 MB
Build✓ · 5s
Tests✗ · 3sran, no count parsed
Known vulns0(pip-audit)
Repo132 files~9,224 lines of source · 24.2 MB · 2 CI workflows · tests dir

Answers from our run

Does laya-mlx build from source?

Dependencies installed in 17 seconds (51 packages), and the build succeeded in 5 seconds. We cloned commit ca5940a into a clean Debian container with 3 CPUs and no project-specific setup.

Do laya-mlx's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does laya-mlx have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use laya-mlx?

Linux, Windows, or Intel Mac deployments: the README requires Apple Silicon and macOS 14 or newer.

What are the alternatives to laya-mlx?

Laya, MLX LM, Jev. Our laya-mlx run installed 51 packages and built in 5 seconds, but its tests stopped in 3 seconds because mlx was unavailable in Debian, confirming that this is an Apple Silicon choice rather than a portable Python runtime.

Setup3/551 packages are manageable, but Mac hardware and model downloads are required
Docs5/5Limits, validation, benchmarks, routing, and attribution are specific
Community4/56,798 stars, an October 2 push, and no open issue or PR
Maturity3/5v0.3.0 is well documented but remains a selective platform port

Who it’s for

Apple Silicon developers who need local typed decisions with no cloud API.
Teams choosing among fixed labels, rubric levels, or true-and-false propositions.
Researchers comparing an MLX port with the pinned upstream Laya implementation.
macOS applications that can keep one or more checkpoints resident in memory.

Who it’s NOT for

Linux, Windows, or Intel Mac deployments: the README requires Apple Silicon and macOS 14 or newer.
Teams expecting a cross-platform test run: our Debian suite stopped during collection because Python could not import mlx.
Users needing training or fine-tuning: the README says those remain in the upstream project.
Services that require upstream batch, long-document, hook, or server APIs: this selective port says it does not include them.
Large taxonomies that must score every label directly: the README warns that choices share a token budget and can lose label distinctions.

Setup reality

Our sandbox installed commit ca5940a in 17 seconds with 51 packages and 150 MB on disk. The build succeeded in 5 seconds. Tests failed after 3 seconds with exit 4 because tests/conftest.py could not import mlx. Pip-audit found 0 known vulnerabilities.

The supported path needs Apple Silicon, Python 3.11 or newer, macOS 14 or newer, and a separately downloaded checkpoint. After that download, inference can stay local without a cloud credential.

Our Debian container was outside the documented target. The log shows ModuleNotFoundError: No module named 'mlx'; it does not show a failed model assertion or how the suite behaves on supported Mac hardware. There is no Dockerfile.

Laya-MLX runs typed decisions only on Apple Silicon

Laya-MLX takes state plus one or more typed questions and returns probabilities for named choices, ordered score levels, or a true proposition. It does not produce prose or JSON token by token. The port implements Laya's ModernBERT encoder, decision transformer, scoring head, and action head in Apple's MLX framework. That gives a Mac application a local decision component without PyTorch, a Transformers runtime, or a cloud request after the model files have been downloaded.

Three checkpoints are supported: a 421M-parameter English model with a 512-token context, a 322M multilingual model with 1,024 tokens, and a 421M typed-decisions model with 1,024 tokens. The project publishes converted FP16 weights on Hugging Face and can convert pinned upstream checkpoints. Training and fine-tuning remain upstream. Laya-MLX is explicit that it is an independent port, not an official Convai Innovations release and not a complete copy of the upstream API.

The 150 MB environment does not include model weights

Our checkout at commit ca5940a contained 132 files, roughly 9,224 lines of source, and occupied 24.2 MB. Installation pulled 51 packages in 17 seconds and used 150 MB on disk. Those numbers stop before a checkpoint download. The repository has 2 CI workflow files and a tests directory, but no Dockerfile. For the intended platform, the README requires Apple Silicon, Python 3.11 or newer, and macOS 14 or newer.

First load downloads the selected checkpoint; later calls can remain local. The Python API accepts text, dictionaries, or conversation lists as state. A router can keep up to 2 models loaded, choose the multilingual checkpoint from language signals, or select a task-specific checkpoint explicitly. That convenience costs memory, especially when models stay resident. Published project figures list peak allocation separately for the English and multilingual checkpoints, but those M3 Max numbers were not produced by our lab.

What happened when we ran it

Our Debian sandbox installed commit ca5940a in 17 seconds and built it successfully in 5 seconds. The container used 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit found 0 known vulnerabilities among the 51 installed packages. Our measurement method covered package installation, build, the repository test command, and audit. It did not download a checkpoint or run inference.

Test collection failed after 3 seconds with exit 4. The traceback ends in tests/conftest.py, where import mlx.core as mx raised ModuleNotFoundError: No module named 'mlx'. The log does not report a failed assertion or any completed test count. Debian is outside the README's supported Apple Silicon macOS target, but the defensible lab finding is simply that a successful install and build did not make the suite runnable in this fresh container.

The M3 Max figures are project benchmarks, not our result

The README reports a 13.42 ms median for one short English question and 7.39 ms for the multilingual checkpoint on an M3 Max with 40 GPU cores and 128 GiB of memory. It says timing includes prompt preparation, tokenization, synchronized inference, calibration, and result formatting, while excluding model load. Batch throughput uses a different 50-question workload. Those boundaries are unusually clear, but the numbers should not be transferred to another Mac or a longer prompt without measurement.

Validation also receives careful treatment. The project says all 3 checkpoints matched upstream selected answers on 63 validation questions in FP16 and FP32, and it publishes raw benchmark artifacts. That is port fidelity on named fixtures, not general accuracy. The runtime even clamps one raw calibration bucket because its stored temperature could make a coin flip appear nearly certain. This is the right kind of warning: a calibrated probability still needs task-specific evaluation before it controls an action.

Shared option tokens limit very large label sets

Choice labels share one token budget. With hundreds of labels, each may receive only a few tokens, causing distinct options to collapse into similar spans. The optional shortlisting helper embeds the state and labels, keeps a top subset, and sends that subset through one decision call. It can make a large taxonomy workable, but the resulting probabilities cover only retained labels. The README says a dedicated bi-encoder usually produces a better shortlist than the decision checkpoint's own encoder.

Release v0.3.0 shipped on October 2, 2026, the same date as the latest repository push. GitHub showed 6,798 stars and 0 open issues or pull requests. The release adds selected upstream behavior while declining full parity, and the README names omitted batch, long-document, hooks, and server APIs. Laya-MLX is a focused Mac runtime with unusually inspectable evidence. Its platform boundary is firm: test it on the exact Apple hardware that will run it, or pick a different implementation.

Alternatives

ProjectWhat it isPick it when
Laya gh↗The upstream open-weight typed-decision implementation and training home.pick this instead when upstream API coverage, training, or non-MLX development matters more than a native Mac runtime.
MLX LMApple's MLX toolkit for running and fine-tuning generative language models.pick this instead when you need text generation on Apple Silicon rather than fixed typed decisions.
JevA hosted typed-decision model accessed through TypeSafe AI and gateway providers.pick this instead when cross-platform API access matters more than local open weights.

What people are saying

  1. [velocity-scout] mizorewww/laya-mlx

Sources

  1. Laya-MLX README
  2. Laya-MLX benchmark methodology
  3. Laya-MLX v0.3.0 release
  4. Upstream Laya project

More ai tools reviews

embodied-jev · underclass · minecraft-agent · laya-coreml · CometixCode · openJev-verdict-2.0 · the whole board →