mrkeyoor.com_
Mon 05 Oct 06:28 UTC
AI Toolsevaluationupdated 05 Oct 2026

Dream-RSI review

Dream-RSI is a research paper and project page that evaluates a new coding-agent exploration method across 8 tasks. It treats the tree of earlier attempts as a replayable world, then tests new branching and stopping policies against recorded outcomes before using a revised policy online.

Verdict

Our lab found no executable ecosystem or Dockerfile at commit 4149ea9, and the README still marks the full codebase plus reproduction scripts as being prepared. Read Dream-RSI for its specific idea of replaying discovery trees to improve exploration policy, not as a tool you can adopt today. Revisit it when the promised artifacts and a clear repository license arrive; until then, OpenEvolve or CodeEvolve is the practical choice for hands-on work.

We ran it

Screenshot of Dream-RSI (github.com/zhengkid/Dream-RSI)

Answers from our run

Did you run Dream-RSI yourself?

No. GitHub reports no primary language for it, and it carries no manifest our lab installs from, and no Dockerfile, so there was nothing standard to install, build or test. This review is written from the repository's own documentation.

Who should not use Dream-RSI?

Developers looking for software they can install today: the README says the full codebase and reproduction scripts are still being prepared.

What are the alternatives to Dream-RSI?

OpenEvolve, CodeEvolve, SkyDiscover. Our lab found no executable ecosystem or Dockerfile at commit 4149ea9, and the README still marks the full codebase plus reproduction scripts as being prepared.

Setup1/5No runnable ecosystem, Dockerfile, or released codebase
Docs3/5Clear README and paper, but no implementation instructions
Community2/51,350 stars and 5 open issues or PRs around a paper release
Maturity1/5No tagged release, code, reproduction scripts, or repo license

Who it’s for

Researchers studying agent exploration, program discovery, and recursive self-improvement at the policy layer.
Teams comparing methods for reusing expensive coding-agent histories without rerunning every evaluator.
Readers willing to study a paper now and wait for the implementation and reproduction artifacts.
Authors of evolutionary coding systems who want a concrete way to separate the coding agent from its exploration controller.

Who it’s NOT for

Developers looking for software they can install today: the README says the full codebase and reproduction scripts are still being prepared.
Evaluators who need to reproduce the paper's reported results now: discovered programs are also marked as being prepared.
Projects that require a clear reuse license: GitHub reports no license for the repository.
Buyers assessing a maintained product release: the repository has no GitHub release, package, Dockerfile, or supported executable ecosystem.
Teams that need a replay policy to test unexplored actions: the paper's simulator reorders and selects outcomes already present in recorded discovery trees, so it does not reveal results for branches that were never collected.

Setup reality

We did not run commit 4149ea9 because the repository had no supported executable ecosystem and no Dockerfile. The README says code and reproduction scripts are being prepared, so there was no honest install, build, or test command for our sandbox to attempt.

What is available is an English README, the paper PDF, diagrams, and an 8-task project summary. The README links an interactive walkthrough, but it does not provide a package, environment file, model credential list, or runnable entry point.

GitHub showed 1,350 stars, but reproducing the work will require more than a coding agent: the method depends on stored discovery trees, evaluators, policy generation, and domain-specific tasks. Exact setup remains unknown until the authors publish the promised artifacts.

1 replay tree changes the cost of policy experiments

Dream-RSI's useful idea is smaller and more concrete than the name suggests. A coding agent explores candidate programs, an evaluator records what happened, and the resulting decisions form a tree. The method then treats that finished tree as a replay environment. Another exploration policy can choose different recorded branches, change their order, group work differently, or stop earlier without asking the coding agent and evaluator to repeat those completed attempts. The object being improved is the exploration controller, while the underlying coding agent stays unchanged.

The paper organizes that loop into 3 stages. An online policy first drives discovery and logs a structured history. Dream-RSI converts accumulated trees into a simulator pool, then generates and evaluates candidate exploration policies inside that pool. A selected policy returns to online discovery, adding more history for a later round. This division is the paper's strongest design choice because it makes branching, concurrency, and stopping explicit. Readers can judge the exploration strategy separately from whatever model proposes the actual code.

Replay cannot score a branch that history never visited

The simulator uses realized code-execution outcomes already stored in a discovery tree. That makes policy comparison cheap in the paper's terms, but it also bounds what replay can reveal. A candidate policy can traverse recorded choices in a new way; it cannot know the outcome of an action absent from the tree. The paper's loop partly addresses that boundary by redeploying improved policies online and collecting fresh trees. Still, the quality and variety of the initial history influence what the offline policy search can learn. This is experience reuse, not a general model of every possible program.

The authors evaluate the framework across 8 tasks in 3 domains: algorithm engineering, mathematical optimization, and GPU kernel engineering. The paper reports improvements or lower discovery cost in several settings and supplies detailed prompts in its appendix. Those are paper claims, not results from MrKeyoor's lab. The repository currently provides no implementation or reproduction scripts with which to check them. For a prospective adopter, the most important result is therefore unavailable: whether the published method can be rebuilt outside the authors' internal environment with the same task machinery.

What happened when we ran it

We did not run commit 4149ea9. The lab found no supported executable ecosystem and no Dockerfile, so there was no install, build, or test target to invoke in the unprivileged sandbox. That matches the README's release table, where the paper and project page are available but the discovered programs, full codebase, and reproduction scripts are marked as being prepared. Calling this a failed build would be misleading. It is a publication repository awaiting the artifacts that would turn the method into testable software.

The repository also has no detected license and no GitHub release. GitHub showed 1,350 stars, 5 open issues and pull requests, and a latest push on 2026-09-16. One open issue asks for the artifacts on Hugging Face, while a pull request proposes listing independent implementations. Another issue asks whether the fixed boundary between the base model, evaluator, and exploration policy is intended as a safety property. These are sensible questions for a young research release, but open discussion is not a substitute for runnable evidence.

The paper is readable before the code exists

The English README gives a compact account of the mechanism, links the full PDF, and distinguishes available items from work still being prepared. The paper explains the replay simulator, policy loop, task families, comparisons, and prompts. That is enough to understand or critique the idea. It is not enough to estimate deployment cost, required model credentials, storage format, evaluator interfaces, or the effort needed to plug in a new domain. Any review claiming easy setup now would be describing an imagined implementation rather than this repository.

A missing license adds another practical stop sign. GitHub reports no repository license, and the paper itself carries a copyright notice. Reading and discussing the work is straightforward; reusing repository assets or an eventual implementation needs terms that are not currently stated here. Researchers should also distinguish the README's aspiration from a commitment with a date. It says the full codebase is being prepared, but it gives no release schedule.

OpenEvolve is the runnable choice today

OpenEvolve and CodeEvolve already provide code for evolutionary program search, so they are better starting points for an experiment this week. SkyDiscover covers a wider scientific and systems-discovery workflow. AI Scientist reaches farther toward end-to-end research automation, which makes it less direct as a comparison but more suitable when paper production is part of the task. Dream-RSI remains the one to read when the question is how to improve an exploration policy from stored discovery histories. The decision can change after a code release; at 4149ea9, there is nothing here to install.

Alternatives

ProjectWhat it isPick it when
OpenEvolveAn available open-source system for evolving programs with language models and evaluators.pick this instead when you need runnable code for evolutionary program search now.
CodeEvolveAn open implementation aimed at algorithm discovery and code optimization.pick this instead when reproducible code evolution matters more than Dream-RSI's replay-policy idea.
SkyDiscoverAn agent framework for scientific, algorithmic, and systems discovery.pick this instead when you want a broader discovery platform with code you can inspect and run.
AI ScientistA research automation system that covers idea generation, experiments, and paper production.pick this instead when the target is an end-to-end research workflow rather than exploration-policy optimization.

What people are saying

  1. [velocity-scout] zhengkid/Dream-RSI
  2. [hackernews] Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Sources

  1. Dream-RSI README and release plan
  2. Dream-RSI paper PDF
  3. Dream-RSI issues and pull requests
  4. Dream-RSI repository metadata

More ai tools reviews

jev-experiments · OrcaBonsai-27B-Uncensored · NanoJev · uplifting-biomolecular-modeling · procedural-film · jev-review · the whole board →