Five video backbones make architecture comparison the point
OpenWAM gives robot-learning researchers one place to swap model architecture, video backbone, action network, dataloader, training behavior, and evaluation client. Its main contribution is experimental control. A lab can change one part through Hydra configuration while keeping the rest of the path recognizable. The repository also publishes OpenWAM-Alpha checkpoints and connects training to deployment through a WebSocket policy server.
The scope is wider than one model. The README documents five supported video backbones, optional visual encoders, a Qwen3-VL understanding component, and adapters for benchmark families including LIBERO, RoboTwin, RoboCasa, VLABench, and EBench. That breadth pays off when your research question compares couplings between video prediction and action learning. If you only need to fine-tune one policy, much of this machinery becomes maintenance.
A successful 11-second build does not make this a small setup
Our checkout held 1,603 files, about 319,960 lines of source, and 35 MB before installation. The Python install succeeded in 90 seconds, pulled 148 packages, and occupied 6,465 MB. The build completed in 11 seconds. Those are manageable repository mechanics on a development box, though the disk footprint arrives before any video backbone, checkpoint, or benchmark dataset.
The native guide asks for Python 3.10 or newer and recommends PyTorch 2.7.1 with CUDA 12.8. DeepSpeed compiles during installation, so gcc must be present. A separate Docker route exists even though the repository scan found no root Dockerfile; it uses project container documentation and a Compose file. The optional Cosmos-Predict2.5 path adds a Git submodule and compiles transformer-engine with nvcc.
What happened when we ran it
Our sandbox installed commit 48bd67b in 90 seconds and built it in 11 seconds. It ran in an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Pip-audit found 0 known vulnerabilities. The test command ran for 424 seconds and exited with code 1.
Pytest reported 1,671 passed, 42 failed, 37 skipped, and 130 collection or setup errors in the supplied 1,843-test run. The log tail shows repeated RoboTwin dataloader cases raising ImportError: libGL.so.1: cannot open shared object file. That tells us the fresh image lacked that shared library. It does not tell us whether every failure and setup error had the same source.
The distinction matters. More than 1,600 passing tests show substantial executable coverage, while the red result means commit 48bd67b cannot enter a strict release pipeline unchanged under this environment. A team evaluating OpenWAM should repeat the suite in its intended CUDA image, retain the full error report, and treat any remaining failures separately. Our sandbox did not train a policy or measure model quality.
Normal training asks for 640 GB of GPU memory
The quick start recommends 8 GPUs with 80 GB of VRAM each for normal training with Wan2.2-5B or smaller backbones. It starts by downloading LIBERO and a video model, then runs a 20-step debug job with checkpoints at steps 10 and 20. The first deployed inference can take longer because compilation warms up. This is a research-compute commitment, even though the shell commands look tidy.
Asset preparation changes configuration files to point at downloaded backbones and benchmark data. Released checkpoints keep their own configuration, while fine-tuning still needs benchmark data and a matching model setup. Reproducibility therefore spans source commit, Python packages, CUDA stack, asset revisions, generated paths, Hydra overrides, and the benchmark environment. The repository documents each part, but your team must pin the resulting whole.
Two open data issues limit exact reproduction
Open issue 42 asks whether the authors will release the filtering pipeline or filtered pretraining data. Issue 41 asks why the listed InternData-A1 frame and hour figures differ from the source dataset paper. Neither question proves that the reported research is wrong. Both matter to a lab trying to reconstruct the training mixture or compare its own run on equal data.
GitHub showed 940 stars, 2 open issues, and 1 open pull request. The repository was pushed on October 1, 2026, and its active pull request concerns exact-step training resume behavior. There is no tagged latest release. The activity is current, but this is a codebase created in September 2026 with research questions still being settled in public. Pin a commit rather than expecting release semantics.
A 6,465 MB environment makes OpenWAM a lab trial
Choose OpenWAM when modular world-action research is the job and the lab already owns the GPU, data, and benchmark burden. Its configuration model and broad adapters can save weeks of stitching together unrelated research repositories. The cost is visible in our 6,465 MB environment, the failed full suite, and the README's 640 GB training recommendation. For a narrow manipulation policy or a small inference service, LeRobot, OpenVLA, or openpi gives you a more focused place to begin.

