ERA: Scaling Scientific Discovery via Semantic Code Mutation and Tree Search

An AI system to help scientists write expert-level empirical software

2026-01-01
Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Julie Jake Garrison, Renee Johnston, Anton Kast, Cory Y. McLean, Peter C. Norgaard, Zahra Shamsi, David Smalling, Julie James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Julie Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Julie Johan Kartiwa, Daniel J. Liebling, Julie Jan-Matthis Lueckmann, Paul Raccuglia, Xuefei Wang, Xuefei Wang, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, Michael P. Brenner
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Empirical Research Assistance (ERA), an AI system that automates the creation of expert-level scientific software for "scorable tasks" using a Large Language Model (LLM) combined with Tree Search (TS). ERA achieved SOTA results in single-cell RNA sequencing (discovering 40 novel methods) and COVID-19 forecasting (outperforming the CDC ensemble), while also excelling in geospatial analysis and numerical integration.

The bottleneck of modern science isn't just the lack of hypotheses; it's the manual labor required to write, test, and optimize the software that validates them. While AI has made strides in generating code, most systems remain "one-shot" wonders or limited to simple competitive programming. Google DeepMind and Google Research have recently unveiled Empirical Research Assistance (ERA)—a system that treats the creation of scientific software as a "scorable task," leveraging Tree Search to navigate the vast space of possible algorithms much like AlphaZero navigates a chessboard.

TL;DR

ERA combines Large Language Models (LLMs) with Tree Search (TS) to iteratively rewrite and improve scientific software. By integrating research ideas from actual scientific literature, it doesn't just "guess" code; it explores sophisticated methodologies. The result? It outperformed humans in complex tasks ranging from single-cell genomics to COVID-19 hospitalization forecasting.


The Problem: The High Cost of Scientific Software

Traditional software for science—modeling weather, predicting disease, or integrating genomic data—takes years of expert labor. This creates two major issues:

  1. Limited Exploration: Researchers often stick to their intuition or standard libraries because trying 1,000 different variants of an algorithm is humanly impossible.
  2. Technical Debt: Most scientific code is a "one-off" that isn't optimized for the specific measurable quality metric of the task at hand.

ERA changes this by defining "Scorable Tasks"—problems where a single quality metric (like Accuracy, MAE, or WIS) can define success—and turning an AI loose on them.


Methodology: Beyond Random Mutations

Unlike Genetic Programming (GP) of the past, which relied on random bit-flips or subtree swaps, ERA uses an LLM to perform semantic mutations. It understands the logic of the code it rewrites.

1. The Tree Search Architecture

ERA uses a PUCT (Predictor + Upper Confidence bound applied to Trees) algorithm.

  • Exploitation: It expands on branches of code that have already shown high scores.
  • Backtracking: If a specific line of optimization (e.g., adding a specific feature) plateaus, it can backtrack and explore a completely different branch.

ERA Schematic Figure 1: Schematic of the ERA algorithm. It takes a scorable task and research ideas, feeds them to an LLM, evaluates in a sandbox, and organizes the results in a tree.

2. Injecting Research Ideas

The "secret sauce" is the Research Idea Injection. Instead of letting the LLM wander aimlessly, researchers feed it summaries of highly-cited papers or use AI Research Agents (like "AI Co-scientist") to find novel strategies. The LLM then implements these ideas as code templates for the Tree Search.


Expert-Level Results

Genomics: Single-Cell RNA Seq

In the OpenProblems v2.0.0 benchmark, ERA integrated complex ideas like BBKNN (Batch Balanced K-Nearest Neighbors) and successfully recombined them with ComBat.

  • The Breakthrough: ERA discovered that using ComBat-corrected PCA embeddings within the BBKNN algorithm (a recombination humans hadn't prioritized) yielded a 14% improvement over standard methods.

ZAPBench results Extended Data: Relative performance of ERA replicates compared to base methods in single-cell genomics.

Epidemiology: COVID-19 Forecasting

Predicting pandemic dynamics is notoriously hard due to noisy, lagged data. ERA was tasked with optimizing the Weighted Interval Score (WIS).

  • Performance: ERA's "Google Retrospective" model achieved an average WIS of 26, significantly beating the CDC’s CovidHub Ensemble (WIS 29), which is literally a robust aggregate of dozens of expert teams.
  • Why it won: It moved beyond simple autoregression by hybridizing epidemiological theory with machine learning, discovering a baseline that balanced seasonal trends with short-term deviations.

Discussion: Why Tree Search Matters

The paper provides a critical insight: ERA vs. Best-of-N. Simply prompting an LLM 1,000 times and taking the best result (Best-of-N) is far less effective than using Tree Search. TS allows for meaningful iteration; the model learns from the errors of previous nodes, whereas Best-of-N is just independent sampling.

ModelMethodBatch integration (higher better)Epidemiology (lower better)
Gemini 3.1 ProBest-of-N0.646192.39
Gemini 3.1 ProERA (TS)0.664172.70

Table: Data confirms that regardless of the underlying LLM, Tree Search (ERA) provides a significant performance delta over raw sampling.


Conclusion: The Precipice of Acceleration

ERA isn't just about writing code; it's about autonomous empirical optimization. While the authors admit limitations—ERA excels at scorable tasks rather than building new physical theories from scratch—the implications for areas like drug discovery, geospatial monitoring, and numerical physics are profound.

By reducing the time to test research ideas from months to hours, ERA represents a fundamental shift in the scientific workflow: from "how do we implement this?" to "how do we score this?".


Key Takeaway: The future of scientific discovery lies in our ability to define scorable metrics for reality and then letting AI-driven tree searches tirelessly iterate on the solution.

Find Similar Papers

Try Our Examples

  • Find recent papers published after 2024 that utilize Tree Search or Monte Carlo Tree Search (MCTS) for autonomous code generation and scientific optimization.
  • What is the current SOTA for single-cell RNA-seq batch integration on the OpenProblems benchmark, and how do they handle "batch effect removal" compared to ERA's BBKNN-TS approach?
  • Are there other systems similar to ERA that successfully combine LLM-based code mutation with external AI-driven literature research agents like Gemini Deep Research or Perplexity?
Contents
ERA: Scaling Scientific Discovery via Semantic Code Mutation and Tree Search
1. TL;DR
2. The Problem: The High Cost of Scientific Software
3. Methodology: Beyond Random Mutations
3.1. 1. The Tree Search Architecture
3.2. 2. Injecting Research Ideas
4. Expert-Level Results
4.1. Genomics: Single-Cell RNA Seq
4.2. Epidemiology: COVID-19 Forecasting
5. Discussion: Why Tree Search Matters
6. Conclusion: The Precipice of Acceleration