ERA: Scaling Scientific Discovery through LLM-Guided Tree Search

An AI system to help scientists write expert-level empirical software

2026-01-01
Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Julie Jake Garrison, Renee Johnston, Anton Kast, Cory Y. McLean, Peter C. Norgaard, Zahra Shamsi, David Smalling, Julie James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Julie Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Julie Johan Kartiwa, Daniel J. Liebling, Julie Jan-Matthis Lueckmann, Paul Raccuglia, Xuefei Wang, Xuefei Wang, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, Michael P. Brenner
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Empirical Research Assistance (ERA), an AI system that combines Large Language Models (LLMs) with Tree Search (TS) to autonomously develop expert-level scientific software. ERA achieved state-of-the-art results across diverse domains, including discovering 40 novel single-cell RNA-seq analysis methods and outperforming the CDC ensemble for COVID-19 hospitalization forecasting.

Executive Summary

TL;DR: Researchers from Google DeepMind and Google Research have unveiled Empirical Research Assistance (ERA), an AI system that doesn't just write code—it conducts research. By coupling Large Language Models (LLMs) with a structured Tree Search (TS), ERA explores thousands of potential software implementations to maximize specific scientific quality metrics. The result? Autonomous software that outperforms human experts and official government ensembles in fields as diverse as genomics, epidemiology, and geospatial analysis.

Strategic Positioning: This work moves beyond "AI for Coding" (like GitHub Copilot) into the realm of Automated Scientist Agents. It marks a transition from one-shot code generation to an iterative, search-based discovery process that can ingest and improve upon the latest academic literature.


The Bottleneck: The "Manual" Science Tax

Why does it take years to develop a new genomic integration tool or a robust forecasting model? The paper argues that scientific software development is currently a "manual search" problem. Scientists try a few ideas, tweak some parameters, and settle on a solution that is "good enough."

However, the space of possible algorithms is effectively infinite. Human designers are limited by:

  1. Bias: Favoring known architectures over novel recombinations.
  2. Speed: The inability to rigorously test hundreds of variations of an idea.
  3. Complexity: The difficulty of integrating ideas from multiple disparate papers into a single, cohesive codebase.

Methodology: The "PUCT" Loop

ERA treats software development as a Scorable Task. If you can define a metric (e.g., Weighted Interval Score for COVID-19 or mIoU for satellite imagery), ERA can optimize for it.

The Search Engine

The core of ERA is a PUCT (Predictor + Upper Confidence bound applied to Trees) algorithm. Unlike simple "Best-of-N" sampling, ERA builds a tree where:

  • Nodes represent specific Python implementations.
  • Edges represent "mutations" or rewrites performed by the LLM.
  • The system balances Exploitation (refining high-scoring code) and Exploration (trying radically new ideas).

Injecting Intelligence

A critical innovation is the injection of Research Ideas. ERA doesn't start from a blank slate; it can be fed PDF summaries of the latest SOTA papers. The LLM then acts as a "translator," turning high-level research concepts into executable code within the search tree.

Overall Architecture of ERA Figure 1: ERA Workflow. Note how research ideas are combined with LLM sampling and sandbox evaluation to drive the search.


Results: Beating the Experts

1. Genomics (scRNA-seq)

Batch integration (removing noise from different lab samples) is a holy grail in single-cell genomics. ERA independently rediscovered components of ComBat and BBKNN, but then went further. By recombining the two, it created a hybrid that outperformed every individual human-designed method on the OpenProblems leaderboard.

2. Epidemiology (COVID-19 Forecasting)

ERA achieved a "Google Retrospective" model that beat the official CDC ensemble. This is no small feat; the ensemble is a robust aggregate of dozens of expert teams. ERA's advantage came from its ability to tirelessly hunt for "needle-in-the-haystack" strategies that hybridize epidemiological theory with modern statistical gradients.

COVID-19 Forecast Results Figure 2: ERA vs. CDC Ensemble. The blue jurisdictions indicate where ERA achieved superior (lower) error rates.


Critical Insight: The Power of Recombination

The most fascinating result from the study is that ERA's "superhuman" performance often came not from inventing entirely new math, but from Hybridization.

For instance, in time-series forecasting, ERA discovered that pairing a "stable" seasonal baseline with a "volatile" machine learning model (like LightGBM) created a solution more robust than either parent. This suggests that the next frontier of science isn't just new data, but the systematic recombination of existing human knowledge at a scale impossible for a single scientist to manage.

Limitations & Ethics

The authors are careful to distinguish between Empirical Optimization and Genuine Theory Discovery. While ERA is phenomenal at optimizing software for a score, it doesn't "understand" the underlying biology in a human sense. Furthermore, the democratization of such powerful engineering tools carries dual-use risks, particularly in sensitive domains like pathogen modeling.

Conclusion

ERA serves as a proof-of-concept for the Autonomous Co-Scientist. By turning software engineering into a search-based optimization problem, it allows researchers to focus on the "What" (the scoring metric and the data) while the AI handles the "How" (the implementation and optimization). As LLMs and inference-time compute continue to scale, the "manual tax" on scientific progress may soon become a relic of the past.


Takeaway: If your scientific problem has a score, ERA can likely find a better way to solve it than a human expert.

Find Similar Papers

Try Our Examples

  • Examine recent papers that utilize Monte Carlo Tree Search (MCTS) or PUCT for automated program synthesis and code optimization in scientific computing.
  • Which study first introduced the concept of using LLMs as mutation operators in Genetic Programming, and how does ERA's approach to "recombination" differ from standard crossover techniques?
  • Investigate the performance of Agentic workflows in the GIFT-Eval benchmark compared to traditional Zero-shot foundation models for long-term time series forecasting.
Contents
ERA: Scaling Scientific Discovery through LLM-Guided Tree Search
1. Executive Summary
2. The Bottleneck: The "Manual" Science Tax
3. Methodology: The "PUCT" Loop
3.1. The Search Engine
3.2. Injecting Intelligence
4. Results: Beating the Experts
4.1. 1. Genomics (scRNA-seq)
4.2. 2. Epidemiology (COVID-19 Forecasting)
5. Critical Insight: The Power of Recombination
6. Limitations & Ethics
7. Conclusion