Jev sees seven actions and structured state, not pixels
TypeSafe Mario gives a model a narrow job: choose Mario's next controller macro. It does not send screenshots. The parser reads emulator telemetry and RAM, then produces JSON covering player motion, jump trajectory, enemies, terrain, response delay, recent control, and episode progress. Jev chooses among seven legal actions, including right_run_jump, left, and noop. That constraint makes the experiment easier to inspect than an agent working from raw video and arbitrary button sequences.
Each request asks Jev for three judgments over the same state. A Choice selects the controller action, a Noul estimates whether a forward jump is useful, and a Score rates immediate danger. Timing arithmetic stays in Python. The parser calculates facts such as whether a jump must begin during the current decision window, while Jev owns the final move. By default, the chosen macro lasts at least 8 emulator frames.
The useful output is the decision record, not a promise that Mario wins
The dashboard shows the current action, probability distribution, confidence, Jev latency, jump probability, danger score, reward, and parsed state beside the game. Every decision also goes into a timestamped JSONL file with the canonical model state, fuller debug state, probabilities, latency, reward, and outcome. That record is the project's best reason to exist. You can inspect what the model saw and compare it with what happened next.
The README does not claim a completion rate, speed record, or trained-policy benchmark, and neither do we. A headless --display none option exists, while the default command opens a recordable dashboard and asks for a decision every 8 emulator steps. Pressing R restarts an episode without closing the dashboard. These details make it suitable for a controlled demo, but they do not establish that Jev reliably clears World 1-1.
What happened when we ran it
Our sandbox installed commit ca22449 in 34 seconds. The install added 48 packages and occupied 71 MB in a fresh Debian container with 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges. The build completed in 19 seconds. Pip-audit found 0 known vulnerabilities in the installed Python environment.
Pytest finished in 9 seconds with 10 passed and 0 failed. The repository has 14 files, about 1,778 lines of source, one CI workflow, and a tests directory. It has no Dockerfile. Those tests cover the state and dashboard modules; the tree has no test file for the runner or policy modules. We did not supply a ROM or TypeSafe credential, so this run says nothing about game completion or API latency.
Setup requires assets and permission the repository cannot provide
The package requires Python 3.13 or newer. Its base dependency is typesafe-sdk; the mario extra adds Gymnasium, gym-super-mario-bros, nes-py, NumPy, and Pygame. The README's setup uses PowerShell commands and asks you to set TYPESAFE_API_KEY. A state-demo command prints the exact JSON and text without launching the game or calling the API, which is a sensible first check before money or game assets enter the loop.
The repository includes no Nintendo ROM. You must bring a lawful local Super Mario Bros. setup, and the README says so plainly. That makes the 71 MB installed footprint only part of the setup. You also need a compatible emulator environment, local game data, and a display for the default dashboard. Headless mode removes the display, not the game file or API requirement.
One failed Jev call can stop the episode
Open issue 1 identifies a sharp runtime failure: runner.py calls the pending request's result() without handling an exception. A timeout or rate-limit error can therefore end the dashboard run. The synchronous headless path calls the policy directly and has the same broad exposure. This matters because the controller depends on a network request during play, where transient failures are ordinary rather than exotic.
Pull request 2 proposes keeping the previous action after a dashboard request fails and retrying on the next eligible frame. It also adds an MIT license. Neither change is merged. On the default branch we reviewed, a failed call can still halt play, and there is no license granting reuse rights. A separate open pull request adds a Makefile for demos, but it does not change those two adoption blockers.
September code and October pull requests show interest, not maturity
GitHub showed 433 stars, 53 forks, and 3 open issues and pull requests when checked. The last code push was September 16, 2026. Issue activity continued through September 23, and the newest pull request opened October 3. There is no published GitHub release, while pyproject.toml reports version 0.1.0. This is a new experiment with visible attention, not a settled package.
TypeSafe Mario is worth reading because its boundary is crisp: Python computes the facts, Jev chooses from seven actions, and JSONL preserves the evidence. The 10 passing tests make that code easier to trust as an example. For reuse in another game or an unattended evaluation service, wait for a license and failure handling, then expect to replace the Mario-specific parser, actions, and goal yourself.

