AI can learn to research better.

Dream-RSI turns experiment history into a replay environment that helps choose better paths for future research.

Independent INEMA guide · Sources checked on Sep 19, 2026
Introductory reading · Original figures by the authors

Original figure: explore with an agent, record the tree, simulate policies, and return to exploration.
The Dream-RSI research cycle. Figure 1, Zheng et al. (2026). Original source ↗

The next experiment is also a decision.

Executing an idea perfectly does not guarantee that it is worth testing. In AI research, deciding where to persist, explore alternatives, or stop can matter as much as writing a good program.

Imagine a map of attempts

Each experiment leaves a path and a result. Later researchers can traverse this map, compare decisions, and build on what has already been measured. The analogy helps explain the method, but the actual map consists of structured execution records.

The strategy is what improves

An exploration policy is the rule that determines how to allocate attempts. In Dream-RSI, this policy is editable code. The coding agent and evaluator remain fixed during the controlled comparison.

The distinction that guides this page: the study demonstrates self-improvement in how search is organized. It does not demonstrate a model rewriting its own weights or an autonomous explosion of intelligence. Method in the paper ↗

One loop builds the map for the next.

The process alternates real work with evaluation over records. The two phases serve different purposes.

1

Explore and record

The current policy guides the agent. Programs are created and evaluated; attempts, relationships, and results form a tree. This phase uses real executions.

2

Turn history into replay

The tree becomes a queryable environment. A policy can visit branches in a different order and choose when to stop exploring them, using results that already exist.

3

Compare candidate policies

An agent develops variations of the exploration code. Replay provides feedback for selecting a policy that balances quality, work, and parallelism.

4

Return to research and expand the history

The selected policy guides the next round. New executions add trees to the simulator set. Improving the strategy and expanding the history complete the cycle.

HistoryReplayPolicyReal experimentNew history

See the mechanism in the authors’ figures.

Images downloaded from the original site and preserved, with explanatory text in Portuguese. Click a figure to open it at full size.

History becomes a simulator. On the left, real attempts build the tree. On the right, other policies traverse parts of that same tree. Replay changes navigation decisions without inventing results for unknown branches.
History becomes a simulator. On the left, real attempts build the tree. On the right, other policies traverse parts of that same tree. Replay changes navigation decisions without inventing results for unknown branches. Figura 2, Zheng et al. (2026). Source ↗
Fewer calls, programs with shorter runtimes. Read the two axes: the horizontal axis counts calls to the discovery agent; the vertical axis measures the average runtime of the resulting programs on test data. Moving down and left is desirable in this chart.
Fewer calls, programs with shorter runtimes. Read the two axes: the horizontal axis counts calls to the discovery agent; the vertical axis measures the average runtime of the resulting programs on test data. Moving down and left is desirable in this chart. Figura 3(b), Zheng et al. (2026). Source ↗

What was measured—and against what.

The paper evaluates eight tasks across three domains. There is no single metric: runtime, agent calls, and solution quality answer different questions.

Lasso · Values reported in Figure 3(a). Lower is better in both numeric columns.
Policy / modelDiscovery callsFinal average runtime (ms)
Fixed · Gemini-3.1-Pro5503,587.1
Dream-RSI · Gemini-3.1-Pro3172,931.0
Fixed · Gemini-3.7-Flash3,2002,516.7
Dream-RSI · Gemini-3.7-Flash1,8792,350.6

The average runtime combines six test datasets. “Compute” in this table means the cumulative number of calls, not dollars, energy, or the total time spent on all research. Model names are kept as published by the authors. Paper, section 4.1 ↗

Compare on the same basis.

Visualization of the published numbers. It does not run AI or estimate a financial bill.

Fixed exploration550 calls
Dream-RSI317 calls
42.4% fewer discovery calls.

Final average runtime: from 3,587.1 to 2,931.0 ms.

The 162× figure needs context

It compares 51,200 SimpleTES calls with 317 Dream-RSI calls on Lasso. The baselines use different models; this is not a controlled comparison that changes only the policy. Do not treat this factor as a general savings promise.

There are also results without a win

In mathematics, Dream-RSI improves the sum-difference result and ties on circle packing among the listed methods. On autocorrelation, SimpleTES retains the best value; Dream-RSI uses a smaller generation budget.

Four KernelBench charts: VGG16 and LayerNorm use fewer generations; ConvDiv and ConvMax achieve higher performance at comparable budgets.
A different domain, different comparisons. VGG16: 2.43× fewer generations; LayerNorm: 1.79×, with comparable performance. ConvDiv and ConvMax: 2.09× and 1.44× higher performance, respectively, at comparable budgets. Figure 4, Zheng et al. Source ↗
ConvDiv evolution: best result improves by round, alongside changes in the number of evaluated attempts.
Effort adapts. In this ConvDiv run, attempts decrease and then increase. This is an observed behavior, not a universal rule that “easy progress = less compute.” Figura 6, Zheng et al. Source ↗

Understanding also means knowing where the idea ends.

Replay cannot see the future.

It answers questions about the space already explored. A branch that was never executed does not gain a reliable result just because a simulation exists. That is why new real rounds remain necessary.

Zero new executions does not mean zero total cost.

The history had to be produced. Generating and reviewing policies, storing records, and running replay also require resources. The reported savings in discovery calls do not capture the full cost.

A policy that performs better on history can fail beyond it.

Including the current policy among the candidates makes it possible to choose one that does not lower the score on that replay. This does not guarantee the same result on new tasks. The learning plan separates development and test trees to make this risk visible.

The agent cannot look up the answer before deciding.

Appendix B restricts the policy to the observed prefix: only nodes already revealed can guide the next decision. Choosing branches by looking at hidden results would leak information and invalidate the comparison.

“Less guidance was better” has a specific scope.

The study’s ablation tests one specific way of adding directional trajectory summaries to prompts. This mechanism hurt results under the evaluated conditions. It is not evidence that all instructions, memory, or guidance are harmful.

AI history is an analogy, not this study’s dataset.

Hinton, GPUs, Transformers, AlphaGo, and AlphaFold help discuss dependencies between discoveries. The Dream-RSI presented here uses trees recorded in discovery tasks; it does not simulate all of scientific history or demonstrate rediscovering the future from a historical cutoff of knowledge.

Status of the work: September 2026 preprint. As of Sep 19, the official repository still announced that the full code was being prepared. This guide is an independent educational analysis and is not affiliated with Google or DeepMind. Check current status ↗

Where AlphaEvolve fits into this story.

AlphaEvolve combines language models and automated evaluation in an evolutionary program search process. Dream-RSI focuses on improving the policy that organizes exploration. Both contribute to automated research, but their mechanisms and results should be presented separately.

DeepMind reports that a heuristic discovered by AlphaEvolve for Borg recovers an average of 0,7% of Google’s global computing resources. The source does not convert this result into “millions of dollars”; that narrator’s estimate will not be presented as fact.

Read DeepMind’s original report ↗

A question to test your understanding.

A policy performed better on every available historical tree. What can we conclude?

Choose an answer to see the explanation.

A project for learning how to evaluate discoveries.

Proposal: a Portuguese-language learning lab that starts with visual reading and progresses to decision experiments. The guide and activity above are available now; the stages below are the expansion plan.

Available
Guide with sources and evidenceExplanation of the cycle, five original figures, an interactive Lasso comparison, limitations, and a comprehension check.
Stage 1
Eight-session sequenceScientific decision → trees → replay → policies → recursion → metrics → limitations → final project. Each session will have a question, an activity, and evidence of learning.
Stage 2
Educational replay labSynthetic trees, a fixed budget, and comparable strategies. Results remain hidden until visited; test scenarios are kept separate. This will be an educational simulation, not a reproduction of the paper.
Stage 3
Critical reading workbookExercises on baselines, quality versus cost, overfitting, leakage, and generalization. Annotated answer keys and a rubric for the final project.
Stage 4
Optional technical reproductionOnce the code is available and infrastructure is defined, evaluate a small task, record the full cost, and publish results, including negative ones.
Who it is for: people interested in AI, educators, and developers who want to understand automated research. The introductory guide requires no programming. The future lab starts without APIs; the technical extension will require Python and basic evaluation knowledge.

Open the full plan Read the source analysis

Go to the source. Check the details.

Primary references support the claims. The video is the editorial starting point; the attached image is not used as scientific evidence.

  1. Original Dream-RSI websiteIntroduction to the method, demo, and source of the five figures on this page.
  2. Zheng et al. — Dream-RSI: Recursive Self-Improvement through Evolving WorldsarXiv:2609.14858v1 · September 14, 2026. Method, tables, ablations, and prompts.
  3. Authors’ repositoryCode status, citation instructions, and release plan.
  4. AlphaEvolve — Google DeepMindSource for the 0.7% result in computing resource management.
  5. Video: “Google is back...”Source of the summary provided for this project; analysis based on the supplied text and comparison with primary sources.

Figures: Zheng et al. / Dream-RSI, kept without overlaid translation. Credits and source URLs are listed in image register. Rights belong to their respective holders. Commentary and activities are INEMA educational content.