mrkeyoor.com_
Sat 03 Oct 07:16 UTC
AI Toolsevaluationupdated 03 Oct 2026

OpenWAM review

OpenWAM is a research stack for training robot policies that learn from video, state, and action data. It gives researchers interchangeable model parts, dataset adapters, training code, deployment tools, and benchmark clients so they can compare world-action model designs without rebuilding the whole pipeline.

Verdict

Our OpenWAM run installed 148 packages and occupied 6,465 MB, then finished with 1,671 passing tests alongside 42 failures and 130 collection or setup errors. That makes it a credible research stack for a robotics lab that can diagnose environment gaps and fund the GPU side, not a ready-made policy server for a small team. Use it to compare architectures and benchmarks; wait if you need a clean suite, lighter hardware demands, or a tagged release.

We ran it

Lab card: what happened when we ran OpenWAMScreenshot of OpenWAM (openwam-official.github.io)
Install✓ · 90s148 packages · 6465 MB
Build✓ · 11s
Tests✗ · 424s1671 passed · 42 failed · 37 skipped · 130 errors of 1843 (pytest)
Known vulns0(pip-audit)
Repo1603 files~319,960 lines of source · 35 MB · 2 CI workflows · tests dir

Answers from our run

Does OpenWAM build from source?

Dependencies installed in 90 seconds (148 packages), and the build succeeded in 11 seconds. We cloned commit 48bd67b into a clean Debian container with 3 CPUs and no project-specific setup.

Do OpenWAM's tests pass?

Not all of them: 1671 of 1843 passed and 42 failed when we ran the project's own test command (pytest), with 130 collection errors. Some failures need services or credentials a bare container does not have.

Does OpenWAM have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use OpenWAM?

Teams that need a green fresh-container test gate: our run ended with 42 failed tests and 130 collection or setup errors.

What are the alternatives to OpenWAM?

LeRobot, OpenVLA, openpi. Our OpenWAM run installed 148 packages and occupied 6,465 MB, then finished with 1,671 passing tests alongside 42 failures and 130 collection or setup errors.

Setup2/56,465 MB install and missing libGL stopped a clean test run
Docs4/5Detailed install, asset, training, deployment, and extension guides
Community3/5940 stars and recent issue activity in a very young repository
Maturity2/5No tagged release and our complete test command failed

Who it’s for

Robotics researchers comparing world-action model architectures under one configuration system.
Labs with CUDA hardware that want to train, fine-tune, deploy, and evaluate robot policies.
Teams using LIBERO, RoboTwin, RoboCasa, VLABench, or the other documented benchmark adapters.
Engineers prepared to manage large model weights, datasets, checkpoints, and benchmark environments.

Who it’s NOT for

Teams that need a green fresh-container test gate: our run ended with 42 failed tests and 130 collection or setup errors.
Labs without serious GPU capacity: the README recommends 8 GPUs with 80 GB of VRAM each for normal training with its 5B or smaller video backbones.
Users seeking a small inference package: installation alone occupied 6,465 MB before model weights and benchmark data.
Researchers who require a published data-filtering pipeline or fully reconciled corpus statistics: open issues 42 and 41 ask for those details.
Production buyers who require tagged releases: GitHub returned no latest release for the repository.

Setup reality

Our fresh Debian sandbox installed commit 48bd67b in 90 seconds, adding 148 packages and using 6,465 MB on disk. The build passed in 11 seconds. Tests failed after 424 seconds: pytest reported 1,671 passed, 42 failed, 37 skipped, and 130 collection or setup errors in the supplied 1,843-test run. Pip-audit found 0 known vulnerabilities.

A useful run needs downloaded video backbones, benchmark data, and usually a released checkpoint. Native setup requires Python 3.10 or newer, a C compiler for DeepSpeed, and the documented CUDA and PyTorch combination. The optional Cosmos route also needs a submodule, CUDA toolkit, and nvcc.

Training is the larger commitment. The README recommends 8 GPUs with 80 GB of VRAM each for its normal 5B-backbone setup. Our test log repeatedly ended on libGL.so.1 missing in RoboTwin dataloader cases, a system library absent from the fresh container.

Five video backbones make architecture comparison the point

OpenWAM gives robot-learning researchers one place to swap model architecture, video backbone, action network, dataloader, training behavior, and evaluation client. Its main contribution is experimental control. A lab can change one part through Hydra configuration while keeping the rest of the path recognizable. The repository also publishes OpenWAM-Alpha checkpoints and connects training to deployment through a WebSocket policy server.

The scope is wider than one model. The README documents five supported video backbones, optional visual encoders, a Qwen3-VL understanding component, and adapters for benchmark families including LIBERO, RoboTwin, RoboCasa, VLABench, and EBench. That breadth pays off when your research question compares couplings between video prediction and action learning. If you only need to fine-tune one policy, much of this machinery becomes maintenance.

A successful 11-second build does not make this a small setup

Our checkout held 1,603 files, about 319,960 lines of source, and 35 MB before installation. The Python install succeeded in 90 seconds, pulled 148 packages, and occupied 6,465 MB. The build completed in 11 seconds. Those are manageable repository mechanics on a development box, though the disk footprint arrives before any video backbone, checkpoint, or benchmark dataset.

The native guide asks for Python 3.10 or newer and recommends PyTorch 2.7.1 with CUDA 12.8. DeepSpeed compiles during installation, so gcc must be present. A separate Docker route exists even though the repository scan found no root Dockerfile; it uses project container documentation and a Compose file. The optional Cosmos-Predict2.5 path adds a Git submodule and compiles transformer-engine with nvcc.

What happened when we ran it

Our sandbox installed commit 48bd67b in 90 seconds and built it in 11 seconds. It ran in an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Pip-audit found 0 known vulnerabilities. The test command ran for 424 seconds and exited with code 1.

Pytest reported 1,671 passed, 42 failed, 37 skipped, and 130 collection or setup errors in the supplied 1,843-test run. The log tail shows repeated RoboTwin dataloader cases raising ImportError: libGL.so.1: cannot open shared object file. That tells us the fresh image lacked that shared library. It does not tell us whether every failure and setup error had the same source.

The distinction matters. More than 1,600 passing tests show substantial executable coverage, while the red result means commit 48bd67b cannot enter a strict release pipeline unchanged under this environment. A team evaluating OpenWAM should repeat the suite in its intended CUDA image, retain the full error report, and treat any remaining failures separately. Our sandbox did not train a policy or measure model quality.

Normal training asks for 640 GB of GPU memory

The quick start recommends 8 GPUs with 80 GB of VRAM each for normal training with Wan2.2-5B or smaller backbones. It starts by downloading LIBERO and a video model, then runs a 20-step debug job with checkpoints at steps 10 and 20. The first deployed inference can take longer because compilation warms up. This is a research-compute commitment, even though the shell commands look tidy.

Asset preparation changes configuration files to point at downloaded backbones and benchmark data. Released checkpoints keep their own configuration, while fine-tuning still needs benchmark data and a matching model setup. Reproducibility therefore spans source commit, Python packages, CUDA stack, asset revisions, generated paths, Hydra overrides, and the benchmark environment. The repository documents each part, but your team must pin the resulting whole.

Two open data issues limit exact reproduction

Open issue 42 asks whether the authors will release the filtering pipeline or filtered pretraining data. Issue 41 asks why the listed InternData-A1 frame and hour figures differ from the source dataset paper. Neither question proves that the reported research is wrong. Both matter to a lab trying to reconstruct the training mixture or compare its own run on equal data.

GitHub showed 940 stars, 2 open issues, and 1 open pull request. The repository was pushed on October 1, 2026, and its active pull request concerns exact-step training resume behavior. There is no tagged latest release. The activity is current, but this is a codebase created in September 2026 with research questions still being settled in public. Pin a commit rather than expecting release semantics.

A 6,465 MB environment makes OpenWAM a lab trial

Choose OpenWAM when modular world-action research is the job and the lab already owns the GPU, data, and benchmark burden. Its configuration model and broad adapters can save weeks of stitching together unrelated research repositories. The cost is visible in our 6,465 MB environment, the failed full suite, and the README's 640 GB training recommendation. For a narrow manipulation policy or a small inference service, LeRobot, OpenVLA, or openpi gives you a more focused place to begin.

Alternatives

ProjectWhat it isPick it when
LeRobot gh↗A robotics toolkit with pretrained policies, datasets, hardware integrations, training, and evaluation.pick this instead when you want a broader route from real robot hardware to policy training and shared datasets.
OpenVLAAn open vision-language-action model and fine-tuning code for robot manipulation.pick this instead when your main goal is adapting one established vision-language-action model rather than comparing modular world-action designs.
openpiA codebase for running and fine-tuning Physical Intelligence robot policy models.pick this instead when you want to evaluate or adapt the released pi model family.

What people are saying

  1. [velocity-scout] OpenWAM-Official/OpenWAM

Sources

  1. OpenWAM repository
  2. OpenWAM README
  3. OpenWAM open issues
  4. Data filtering issue
  5. InternData-A1 statistics issue

More ai tools reviews

GPT-as-Policy · NeuralScreen · mural · recurrent-looped-tranformer · gpu-time · xialingguo-ip · the whole board →