mrkeyoor.com_
Sat 03 Oct 07:17 UTC
AI Toolsevaluationupdated 03 Oct 2026

GPT-as-Policy review

GPT-as-Policy is an English and Chinese research release that compares GPT 6 Astra acting directly as a robot policy with a hybrid policy that can correct actions proposed by a vision-language-action model. It packages evaluation code, selected results, videos, and a rebuildable report, while leaving simulator assets, model checkpoints, and original operations data to their upstream or private environments.

Verdict

Our GPT-as-Policy run installed 260 packages, built in 10 seconds, and then passed 0 tests, so the public report is easy to rebuild but its green test line proves nothing about the robotics code. Use the repository to inspect a well-documented 50-case comparison or to reuse its evaluation scaffolding inside an existing robotics lab. Skip it if you need a self-contained simulator, full experiment reproduction, or published control-loop latency.

We ran it

Lab card: what happened when we ran GPT-as-PolicyScreenshot of GPT-as-Policy (anonymous-report-421.github.io/public-website/?view=1)
Install✓ · 27s260 packages · 147 MB
Build✓ · 10s
Tests✓ · 8s0 passed · 0 failed of 0 (node:test)
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo850 files~55,671 lines of source · 91.4 MB · 0 CI workflows

Answers from our run

Does GPT-as-Policy build from source?

Dependencies installed in 27 seconds (260 packages), and the build succeeded in 10 seconds. We cloned commit 8f3d362 into a clean Debian container with 3 CPUs and no project-specific setup.

Do GPT-as-Policy's tests pass?

Yes: 0 of 0 passed when we ran the project's own test command (node:test). Some failures need services or credentials a bare container does not have.

Does GPT-as-Policy have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use GPT-as-Policy?

Teams seeking a turnkey robotics stack: the README says simulator assets, model checkpoints, raw sessions, and original operations logs are not included.

What are the alternatives to GPT-as-Policy?

OpenPI, Isaac Lab. Our GPT-as-Policy run installed 260 packages, built in 10 seconds, and then passed 0 tests, so the public report is easy to rebuild but its green test line proves nothing about the robotics code.

Setup3/5Report build passed, but new evaluations need several external systems
Docs5/5Clear rebuild, evaluation, security, provenance, and limit notes
Community2/5569 stars, two unanswered issues, and no pull requests
Maturity3/5Strong release artifacts, but no CI and the npm suite found 0 tests

Who it’s for

Robotics researchers studying language models as direct or supervisory control policies.
Reviewers who want selected case metadata, scores, seeds, provenance, and videos beside the paper's claims.
Frontend developers who need to rebuild or inspect the bilingual static report without running a robot simulator.
Labs that already operate Isaac Sim, RoboDojo or RoboLab, OpenPI/JAX, and authorized GPT 6 Astra access.

Who it’s NOT for

Teams seeking a turnkey robotics stack: the README says simulator assets, model checkpoints, raw sessions, and original operations logs are not included.
Buyers who require the full 120-task RoboLab suite: the report covers 10 tasks, and open issue 6 asks whether the remaining scope will be evaluated.
Anyone budgeting robot throughput from published latency data: open issue 5 asks for episode runtime and control-loop latency, and it has no response.
Release gates that treat a green npm test command as coverage: our node:test run passed in 8 seconds after collecting 0 tests.
Operators expecting a container recipe or visible CI checks: our scan found no Dockerfile and 0 CI workflow files.

Setup reality

Our commit 8f3d362 run installed 260 npm packages in 27 seconds and used 147 MB, then built the report app in 10 seconds. The 8-second node:test step exited successfully with 0 passed and 0 failed of 0, and npm audit reported 0 known vulnerabilities.

The report preview and frontend build need no model credential. A new evaluation needs separate simulator and OpenPI/JAX environments, Isaac Sim 5.1 for RoboDojo, a compatible checkpoint, Codex CLI, and authorized access to the fixed gpt-6-astra model at xhigh reasoning.

The app lives under hybrid_rollout/report_site/app. The README says Node 24 was used, although our Node 22 image completed the build. There is no Dockerfile, CI workflow, or tests directory, and the public snapshot omits assets and records needed for full experiment reproduction.

The 37-second install and build path rebuilds the report

Our sandbox installed 260 npm packages in 27 seconds and built the Vite report in 10 seconds. That is the friendly part of GPT-as-Policy. The app under hybrid_rollout/report_site/app turns the preserved report source into one self-contained HTML file, while scores, media, licenses, and citation data sit beside it. You can also serve the prebuilt report with Python and switch between English and Chinese.

That success says nothing about launching a robot. The public repository is a research release around two evaluation paths: GPT 6 Astra Direct chooses actions itself, while the hybrid policy reviews a fresh proposal from π₀.₅ and may correct it. The release includes code, selected results, and provenance, but it deliberately excludes model checkpoints, simulator assets, raw model sessions, private annotations, and original operations logs.

The hybrid policy reports 48% success across 50 cases

The project README reports 10 RoboDojo tasks with five aligned cases per task. Its hybrid policy reaches a mean score of 62.60 and 48% success, with GPT corrections on 14.4% of executed control steps. GPT 6 Astra Direct reaches 26% success, and its mean score is 37.81 over the 48 episodes that had native scores. Both success rates use all 50 selected cases.

Those figures support a narrow conclusion: on this selected evaluation, reviewing π₀.₅ proposals beat direct GPT control. They do not establish performance across the full RoboLab catalogue. The README also says official-model comparisons are reweighted public references rather than same-seed reruns. Open issue 6 asks about extending the work from 10 tasks to the full 120-task suite, and no answer is posted.

What happened when we ran it

We cloned commit 8f3d362 into an unprivileged Debian container with 3 CPUs, 8 GB of RAM, no secrets, and a Node 22 image. The checkout contained 850 files, about 55,671 lines of source, and occupied 91.4 MB. The project app lives several directories below the root, which matters if you expect npm ci at the checkout top level.

Our npm install succeeded in 27 seconds, adding 260 packages and consuming 147 MB. The build completed in 10 seconds. Npm audit found 0 known vulnerabilities across critical, high, moderate, and low severity. These measurements cover the report app only; we made no simulator, policy-model, or paid API call.

The test command exited successfully in 8 seconds, yet node:test reported 0 passed and 0 failed of 0. The repository scan also found 0 CI workflow files, no Dockerfile, and no tests directory. A zero exit code here confirms that the command returned cleanly. It does not show that any assertion ran.

The npm test script points to a directory that is absent

The app's package script calls node --test tests/*.test.mjs, while the measured checkout has no tests directory. That matches the 0-test result. The wider source tree does contain Python test modules, and the README gives separate unittest and pytest commands, but our lab did not run those suites. Treat the npm result as an empty frontend test target, not as coverage for the controller or report.

This distinction is easy to miss because the build performs a protected-runtime verification before Vite. That integrity check can prove that preserved files match expected digests. It cannot replace behavioral tests. A team reusing the frontend should wire its actual tests into the package script and CI before relying on green automation.

A new rollout needs Isaac Sim 5.1 and authorized model access

The public snapshot is not a standalone simulator distribution. RoboDojo requires Isaac Sim 5.1, matching RoboDojo source and assets, and the π₀.₅ OpenPI/JAX checkpoint. The model is fixed to gpt-6-astra with xhigh reasoning, with no provider fallback promised. You must also configure paths, container identity, credential locations, a gateway, and the exact evaluation case.

The missing material is documented rather than hidden. Source revisions and checkpoint identity are recorded, public results align task, scene, and seed, and the release notes explain that deployment paths and gateway endpoints were replaced with generic examples. Still, full reproduction requires upstream assets plus authorized and potentially paid services outside the checkout. A successful 10-second report build should never be presented as a reproduced robotics experiment.

Two open issues expose the remaining decision gaps

GitHub showed 569 stars, two open issues, no open pull requests, and a last push on September 16, 2026. The source release was recent, so that date does not suggest abandonment. Both issues are relevant, though neither had a reply when checked: issue 5 asks for episode runtime and control-loop latency, while issue 6 asks about all 120 RoboLab tasks.

No GitHub release tag exists, despite a checked-in source-release note and an initial public-release commit. For reading the report, commit 8f3d362 is a clear pin. For adopting the controller, wait for your own simulator run before making cost or latency promises. GPT-as-Policy is most useful as a transparent evaluation artifact: its strongest product is the boundary it draws between rebuilding the report and reproducing the experiment.

Alternatives

ProjectWhat it isPick it when
OpenPIThe open repository for π models, policies, checkpoints, and robot-policy workflows.pick this instead when you need to train or run the underlying vision-language-action policy rather than inspect this GPT comparison.
Isaac LabA robot-learning framework built on Isaac Sim with tasks, environments, and training workflows.pick this instead when your first need is a simulation and training platform rather than a fixed evaluation release.

What people are saying

  1. [velocity-scout] anonymous-report-421/GPT-as-Policy

Sources

  1. GPT-as-Policy README
  2. Report app package scripts
  3. Source release notes
  4. Issue 5: runtime and control-loop latency
  5. Issue 6: full RoboLab evaluation
  6. Commit 8f3d362

More ai tools reviews

NeuralScreen · mural · recurrent-looped-tranformer · gpu-time · OpenWAM · xialingguo-ip · the whole board →