It scores supplied actions instead of writing an answer
NanoJev takes a state, a question, and a set of candidates, then returns probabilities over those candidates. Choice accepts 2 to 255 options. Boolean returns a probability for a proposition, while Score spreads probability across 2 to 10 ordered levels and calculates an expected value. That interface suits routing, ranking, and game actions where the valid moves are already known. It does not supply the open-ended text that a chat model would.
The backbone is Qwen3-0.6B with separate decision heads. Candidate paths pass through the same model, and a set-attention head handles Choice questions. The repository's stronger idea is the contract around that model: independent questions share a state but cannot consume one another's answers. A developer can inspect complete distributions instead of parsing a generated label and hoping the model followed a JSON instruction.
The released checkpoint proves four games, not general judgment
The public unified-games-v1 checkpoint covers Maze, Snake, ViZDoom Basic, and ViZDoom Predict Position. The README reports 18,760 questions per data variant and a 274-case held-out game set. NanoJev recorded 128 successes in 128 Basic cases and 27 in 128 Predict Position cases. Those are project-published results, not measurements from our sandbox, and the repository provides recordings plus replay checks for examining them.
Scope matters more than the winning rows in that table. The same checkpoint learned four related environments with known actions and carefully prepared state text. NanoJev reports 4 of 10 Maze successes, below Jev's 7 of 10, even while beating the listed baselines in Basic and Predict Position. The roadmap describes broader long-horizon work as unfinished. Treat the games as evidence that the architecture can learn these tasks, not as proof that it can judge your support tickets or trading rules.
What happened when we ran it
Our sandbox cloned commit 76fdfc9, a 520-file repository with about 39,668 lines of source and a 245.8 MB checkout. npm install succeeded in 15 seconds, added 20 packages, and occupied 25 MB. The audit reported 0 known vulnerabilities across critical, high, moderate, and low severity. The container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges.
There was no npm build script, so we skipped the build rather than substituting an undocumented command. There was also no npm test script, so no tests ran. Our scan found 0 CI workflow files, no Dockerfile, and no tests directory. This result verifies that the small Node dependency tree installs cleanly. It says nothing about checkpoint accuracy, Python dependency compatibility, GPU inference, or the replay claims.
A 15-second npm install is only the side entrance
The README's model quick start switches ecosystems. It asks for Python packages, then downloads a model snapshot and a separate dataset snapshot from Hugging Face. The release document lists 149 model files totaling 4,826,217,398 bytes and 214 dataset files totaling 880,108,238 bytes. Both are public, so credentials are optional, but disk, download time, and model-loading memory belong in any honest setup estimate. None was exercised by our npm run.
Serving also assumes CUDA in the documented command. The pinned toy requirements name Python 3.14.4, PyTorch 2.14.0, Transformers 5.17.0, and an A100 80GB as the environment used by the author. That is unusually specific and useful for reproduction, but it is a demanding baseline. Recorded browser demos are easier: Python's built-in HTTP server can expose the supplied web files without loading the model.
Local responses still differ from the TypeSafe contract
The contract audit is candid about compatibility. Local inference calls the yes-or-no primitive boolean rather than TypeSafe's noul. Structured state objects become Python string representations instead of canonical JSON. Local questions require nonempty text where the upstream interface accepts structured or null descriptions, and local Score responses omit the legend and confidence fields. These gaps matter if existing TypeSafe clients are supposed to switch endpoints unchanged.
They also limit what training data can express. A normalized distribution is useful, but it does not establish that 0.8 corresponds to an 80 percent empirical success rate. The docs make that distinction and separate atomic geometry questions from whole-map planning. That restraint is a reason to study the repository: its own compatibility table tells you where the replica stops.
The September 21 code is active, with 7 open pull requests
GitHub showed 2,492 stars, an MIT license, and 8 combined open issues and pull requests on October 5, 2026. Seven of those 8 items are pull requests, including work on Apple Silicon inference, shared-prefix candidate encoding, verification, extra datasets, and path handling. The last repository push was September 21, while one path-fix pull request arrived September 27. That is recent participation, though none of it has produced a tagged GitHub release.
NanoJev is easiest to recommend as an inspectable research artifact. The 4-game checkpoint, public data, and explicit contract boundaries give you something concrete to test. Production adoption asks for more: run the Python path on your hardware, measure calibration on your own decisions, and decide whether a fixed candidate set fits the product before downloading roughly 5.7 GB of published assets.

