Three MuJoCo tasks make the workbench concrete
EmbodiedJev puts a Franka Panda arm in three simulated jobs: move a red block into a tray, stack it on another block, and carry it over an obstacle. MuJoCo supplies contact and motion, while separate physical conditions decide whether a run succeeded. You can pause, step, reset, replay a trajectory, and export the experiment as JSON from a browser interface. No robot hardware is required.
The project is easier to understand than many embodied-agent demos because it names what the model does and what code still owns. A model selects a stage, skill, or short XYZ and gripper action, depending on mode. The controller converts that choice into inverse kinematics, runs safety checks, moves the arm, and observes the next state. Candidate generation, geometry, and collision handling remain program logic. A successful simulation therefore does not prove the model learned manipulation.
Rule, Jev, chat, and local-model runs share one interface
The rule baseline works without an API key, GPU, or model weights. Other adapters cover TypeSafe Jev, OpenAI-compatible chat endpoints, Anthropic's native API, a structured decision endpoint, and local MiniCPM. The interface can run two or three configurations against the same task and seed, then align their recorded trajectories by simulation time. That is a useful way to inspect choices without confusing replay speed with inference latency.
Input modes matter. Jev receives structured text, which may contain simulator coordinates or coordinates estimated by the local RGB-D detector. Compatible vision models can instead receive raw RGB frames from the external camera, wrist camera, or both, along with robot state. The README keeps those protocols separate and warns against treating them as one score. This is the right habit for a workbench where a small input change can alter what problem the model is solving.
What happened when we ran it
Our sandbox installed commit f08de2e in 24 seconds. Npm added 26 packages, and the installed environment occupied 148 MB. The Vite front end built successfully in 10 seconds. Npm audit found 0 known vulnerabilities. The repository checkout was much larger than its JavaScript dependency tree at 109.9 MB, partly because it includes media, experiment material, Python code, and site assets.
We did not run tests because package.json has no test script or equivalent default target. It does define test:ui and test:site, and the repository contains Python, browser, and site test directories plus two GitHub Actions workflow files. Those signals are better than having no test material. They do not change the lab result: the standard npm test step was unavailable, so our build has no accompanying test count.
That distinction belongs in the buying decision. A 10-second front-end build proves the bundled web assets compile at commit f08de2e. It does not exercise MuJoCo physics, API adapters, credential persistence, Playwright flows, or Python behavior. A maintainer can run the named suites separately, but a contributor looking for one obvious release gate will not find it in package.json.
The web build is only half of local setup
The complete quick start asks for Python 3.11 or newer, Node.js 22.12 or newer, Git, a virtual environment, an editable Python install, npm ci, and npm run build. The Python package then serves the compiled interface on 127.0.0.1:8090. The rule baseline is the sensible first check because it avoids provider access and separates simulator problems from model problems.
MiniCPM is a heavier branch. Its documented route adds Torch, Transformers, and Accelerate, then downloads about 5 GB of weights before allowing for runtime memory. CUDA, Apple MPS, and CPU paths use different numeric formats. Cloud routes avoid that download but need provider credentials and may incur charges on connection tests and experiments. TypeSafe access is invitation-based according to the README.
Open issue 2 reports a JSON Not Found response at the expected localhost root. The thread has a maintainer reply, but the issue remains open. That does not show the current commit fails for everyone, and our lab did not start the service. It does show why setup verification should include the actual browser route after both the Python package and front-end assets are installed.
Cloud vision sends camera frames beyond the machine
By default the service binds to localhost, and saved credentials use the operating system keyring. If secure storage is unavailable, the UI falls back to session-only handling instead of writing a plaintext key. Keys are excluded from browser storage, presets, exports, and logs. A memory-only server mode is available for users who do not want persistent connection profiles.
Local binding does not make cloud experiments local. Structured providers receive task state, and direct-image modes send the enabled cameras' RGB frames plus robot feedback to the chosen service. Connection tests make real requests, while saving a profile does not. The UI's separation between saved and verified connections is good, but the operator still owns provider retention policy, cost, and whether a lab scene is suitable to transmit.
Four open items and no tag keep this in the lab
GitHub showed 257 stars and 4 open issues and pull requests on October 7, 2026. The last push was September 22, and there is no published release. The open queue includes a pull request and issue about preserving unknown token usage instead of recording it as zero. That matters for comparison charts because missing accounting can otherwise look like free inference.
The MIT license, two workflows, detailed limitations, and explicit rule baseline make EmbodiedJev worth trying as a learning and experiment tool. Its boundaries are equally clear: fixed simulated tasks, Chinese documentation, no hardware proof, no tagged release, and no default npm test command in our run. The best first use is a local rule-baseline session, followed by one pinned model configuration whose input and costs you can inspect.

