Laya returns typed decisions instead of prose
A support ticket arrives, and the application needs three answers: which queue gets it, how urgent it is, and whether the customer may leave. Laya expresses those as choice, score, and yes/no questions, then returns probabilities and routing details. It is designed for that narrow job. The output can feed code directly without asking a chat model to follow a JSON format.
The Router chooses among English, multilingual, and task-specific checkpoints. The README also documents batch calls, schema-driven decisions, confidence gates, presets, hooks, and long-document scanning. Integrations cover an HTTP service, MCP, LangChain, LlamaIndex, CrewAI, TypeScript, Java, and .NET. Python 3.10 or newer is the main path, and the default package depends on PyTorch, Transformers, Safetensors, Hugging Face Hub, and NumPy.
The base checkpoints need your own evaluation set
Laya's most useful sentence appears deep in its own limits section: fine-tuning supplies the value on its typed-decisions benchmark. The README says the base English and multilingual checkpoints sit below the majority-class baseline there. That makes Laya a promising decision architecture and training starting point, rather than a universal zero-shot judge you can safely place in front of refunds, bans, or production changes.
The failure modes are unusually well documented. Choice labels can sway the result, negated cancellation requests have produced wrong actions, ordinal score questions are the weakest primitive, and the multilingual checkpoint has an open position-bias issue for score options. Confidence helps rank uncertain answers, but it does not prove correctness. Build a held-out set from the exact language, wording, option order, and consequences your application uses.
Fine-tuning now has a laya-train command, CSV and expected-result loaders, calibration reporting, and export paths. That is useful plumbing, though open issue 963 says the fixed 4-epoch default can collapse toward the class prior on small datasets. A team still needs to split training from evaluation, inspect per-slice errors, select an abstention threshold, and keep a person in the path when a wrong decision is expensive.
What happened when we ran it
Our fresh Debian sandbox installed commit 8a6e132 in 70 seconds. Pip pulled 80 packages, and the environment occupied 5,472 MB on disk. The build succeeded in 6 seconds. Pip-audit found 0 known vulnerabilities. The checkout itself contained 1,023 files, about 121,224 lines of source, and 16 MB before installation.
The test command did not finish cleanly. After 23 seconds, pytest exited with code 3 and reported 2 collection/setup errors out of 2 in the lab summary. Its own final lines also said 64 passed, 0 failed before 2 errors in 10.87s. Both errors came from collection reaching research/eval/test_laya_eval.py, where a module-level sys.exit(1 if FAIL else 0) raised SystemExit: 0.
That log supports a precise conclusion: many checks passed, but the repository-wide pytest command was red because collection encountered a successful SystemExit. We cannot tell from the supplied tail whether the file is intended to run only as a standalone evaluation script. The failed command is still a release-gate problem for anyone who expects plain pytest to be authoritative.
A 5,472 MB install is only the first operating cost
The first prediction downloads a checkpoint, and Router(preload=True) can load all three. An operator must decide which models stay resident, how model downloads are pinned and cached, and whether CPU, CUDA, Apple Silicon, ONNX, or the optional TileLang route matches production. Long input also needs an explicit length setting because the multilingual checkpoint defaults to a 1,024-token limit.
Serving adds another layer. The HTTP extra brings FastAPI and Uvicorn, while MCP is separate and can either load locally or point to a shared Laya server. A remote MCP client loses the shortlist tool according to the README. API keys, device settings, preload lists, thread limits, and checkpoint revisions therefore belong in deployment review, even though our 6-second build needed no secrets.
v0.3.28 is active, with beta-grade movement
GitHub showed 31,011 stars, 47 open issues, and 36 open pull requests on October 6, 2026. The repository was pushed on October 5, and v0.3.28 shipped that day. Its release fixed missing inference backends in the previous three wheels, added training commands, changed the Docker base, and repaired behavior across MCP, ONNX, TypeScript, Java, and calibration code.
That pace is evidence of maintenance and also a warning about churn. The package metadata calls Laya beta software, which fits what we found: thoughtful documentation, broad runtime work, and active fixes around packaging and numerical behavior. Use it for a measured decision service where you own the test set. If you need a small conventional classifier, SetFit or fastText asks you to operate far less machinery.

