ERA: Scaling Scientific Discovery via LLM-Driven Tree Search

An AI system to help scientists write expert-level empirical software

2026-01-01
Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Julie Jake Garrison, Renee Johnston, Anton Kast, Cory Y. McLean, Peter C. Norgaard, Zahra Shamsi, David Smalling, Julie James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Julie Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Julie Johan Kartiwa, Daniel J. Liebling, Julie Jan-Matthis Lueckmann, Paul Raccuglia, Xuefei Wang, Xuefei Wang, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, Michael P. Brenner
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Empirical Research Assistance (ERA), an AI system that automates the creation of expert-level scientific software by combining Large Language Models (LLMs) with Tree Search (TS). It focuses on "scorable tasks" across diverse disciplines, including genomics, epidemiology, and time-series forecasting, consistently matching or exceeding human SOTA performance.

TL;DR

Researchers from Google DeepMind and Google Research have unveiled Empirical Research Assistance (ERA), a system that automates the development of scientific software. By framing software engineering as a "scorable task" and navigating the solution space with Tree Search and LLMs, ERA has effectively surpassed human experts in fields as diverse as single-cell genomics and COVID-19 epidemiology.

The Bottleneck: The "Slow" Cycle of Science

Experimental science is frequently held back by a hidden debt: the manual labor of writing computational pipelines. Whether it's integrating batch effects in genomics or forecasting disease spread, the software is often built on intuition. Consequently, the space of possible algorithms is rarely explored exhaustively. ERA changes the game by treating the creation of this software as an optimization problem where the objective is to maximize a specific quality score.

Methodology: Tree Search Meets Scientific Literature

At its heart, ERA is not just a chatbot writing code; it is a search engine for logic.

  1. Architecture: The system uses a PUCT (Predictor + Upper Confidence bound applied to Trees) algorithm. Each node in the tree is a functional Python script.
  2. The LLM as Mutator: Unlike traditional Genetic Programming that uses random bit-flips, ERA uses an LLM (like Gemini) to read the code of a parent node, understand its performance logs, and rewrite it according to specialized "research ideas."
  3. Knowledge Injection: Research ideas are pulled from PDFs of highly-cited papers or generated by AI synthesis tools. ERA can then "recombine" these ideas—for instance, taking the stability of one model and the feature engineering of another to create a superior hybrid.

ERA Algorithm Schematic

Key Results: Beating the Human Leaderboards

1. Single-Cell Genomics (scRNA-seq)

The task was to remove "batch effects" (technical noise) in single-cell data while preserving biological signals. ERA's generated methods (like BBKNN (TS)) outperformed 300 existing tools. By recombining linear correction (ComBat) with graph-based neighbors (BBKNN), it achieved a 14% improvement over the previous SOTA.

2. Epidemiology: COVID-19 Forecasting

ERA created a "Google Retrospective" model for CDC hospitalization data. It achieved an average Weighted Interval Score (WIS) of 26, significantly better than the official CDC Ensemble's 29. The breakthrough was ERA's ability to discover that seasonal foundations (CMU-climate) hybridized with epidemiological trends (Rtrend) yielded models far more responsive than individual human efforts.

COVID-19 Forecast Performance

3. General Time Series (GIFT-Eval)

In a unified solution search, ERA built a general-purpose forecasting library from scratch using only numpy and pandas. This AI-written library outperformed established foundation models and deep learning models on the GIFT-Eval benchmark.

Deep Insight: The Value of "Needle-in-a-Haystack" Search

The core takeaway is that ERA identifies "needle-in-a-haystack" solutions that humans miss because we are too slow to try subtle variations. The "Breakthrough Plots" in the paper show that scores often plateau before a specific LLM-driven mutation triggers an abrupt jump in quality—a process mirroring human "Eureka" moments but occurring at a million-fold speed.

Breakthrough Plot Example

Conclusion and Future Outlook

While ERA excels at empirical optimization (maximizing scores), the authors note that genuine scientific discovery—reasoning about causal mechanisms—remains a frontier. However, for any scientific task that can be defined by a clear metric, ERA suggests that the bottleneck has shifted from "how to write the software" to "how to define the right score."

Industry Takeaway: For R&D teams, this implies that the future of competitive science lies in Inference-Time Compute. By spending more compute on searching for the best algorithm, rather than just using a pre-trained model, we can achieve superhuman results today.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Upper Confidence bound applied to Trees (UCT) or MCTS for optimizing non-game-playing code generation tasks.
  • Which paper originally proposed the concept of "Genetic Programming" for code evolution, and how does ERA's use of LLMs as a mutation operator fundamentally differ from traditional crossover methods?
  • Find research papers exploring the application of AI agents to cross-domain scientific benchmarks beyond scRNA-seq and epidemiology, specifically in physical chemistry or material science.
Contents
ERA: Scaling Scientific Discovery via LLM-Driven Tree Search
1. TL;DR
2. The Bottleneck: The "Slow" Cycle of Science
3. Methodology: Tree Search Meets Scientific Literature
4. Key Results: Beating the Human Leaderboards
4.1. 1. Single-Cell Genomics (scRNA-seq)
4.2. 2. Epidemiology: COVID-19 Forecasting
4.3. 3. General Time Series (GIFT-Eval)
5. Deep Insight: The Value of "Needle-in-a-Haystack" Search
6. Conclusion and Future Outlook