ERA: Scaling Scientific Discovery via Semantic Code Mutation and Tree Search
An AI system to help scientists write expert-level empirical software
The paper introduces Empirical Research Assistance (ERA), an AI system that automates the creation of expert-level scientific software for "scorable tasks" using a Large Language Model (LLM) combined with Tree Search (TS). ERA achieved SOTA results in single-cell RNA sequencing (discovering 40 novel methods) and COVID-19 forecasting (outperforming the CDC ensemble), while also excelling in geospatial analysis and numerical integration.
The bottleneck of modern science isn't just the lack of hypotheses; it's the manual labor required to write, test, and optimize the software that validates them. While AI has made strides in generating code, most systems remain "one-shot" wonders or limited to simple competitive programming. Google DeepMind and Google Research have recently unveiled Empirical Research Assistance (ERA)—a system that treats the creation of scientific software as a "scorable task," leveraging Tree Search to navigate the vast space of possible algorithms much like AlphaZero navigates a chessboard.
TL;DR
ERA combines Large Language Models (LLMs) with Tree Search (TS) to iteratively rewrite and improve scientific software. By integrating research ideas from actual scientific literature, it doesn't just "guess" code; it explores sophisticated methodologies. The result? It outperformed humans in complex tasks ranging from single-cell genomics to COVID-19 hospitalization forecasting.
The Problem: The High Cost of Scientific Software
Traditional software for science—modeling weather, predicting disease, or integrating genomic data—takes years of expert labor. This creates two major issues:
- Limited Exploration: Researchers often stick to their intuition or standard libraries because trying 1,000 different variants of an algorithm is humanly impossible.
- Technical Debt: Most scientific code is a "one-off" that isn't optimized for the specific measurable quality metric of the task at hand.
ERA changes this by defining "Scorable Tasks"—problems where a single quality metric (like Accuracy, MAE, or WIS) can define success—and turning an AI loose on them.
Methodology: Beyond Random Mutations
Unlike Genetic Programming (GP) of the past, which relied on random bit-flips or subtree swaps, ERA uses an LLM to perform semantic mutations. It understands the logic of the code it rewrites.
1. The Tree Search Architecture
ERA uses a PUCT (Predictor + Upper Confidence bound applied to Trees) algorithm.
- Exploitation: It expands on branches of code that have already shown high scores.
- Backtracking: If a specific line of optimization (e.g., adding a specific feature) plateaus, it can backtrack and explore a completely different branch.
Figure 1: Schematic of the ERA algorithm. It takes a scorable task and research ideas, feeds them to an LLM, evaluates in a sandbox, and organizes the results in a tree.
2. Injecting Research Ideas
The "secret sauce" is the Research Idea Injection. Instead of letting the LLM wander aimlessly, researchers feed it summaries of highly-cited papers or use AI Research Agents (like "AI Co-scientist") to find novel strategies. The LLM then implements these ideas as code templates for the Tree Search.
Expert-Level Results
Genomics: Single-Cell RNA Seq
In the OpenProblems v2.0.0 benchmark, ERA integrated complex ideas like BBKNN (Batch Balanced K-Nearest Neighbors) and successfully recombined them with ComBat.
- The Breakthrough: ERA discovered that using ComBat-corrected PCA embeddings within the BBKNN algorithm (a recombination humans hadn't prioritized) yielded a 14% improvement over standard methods.
Extended Data: Relative performance of ERA replicates compared to base methods in single-cell genomics.
Epidemiology: COVID-19 Forecasting
Predicting pandemic dynamics is notoriously hard due to noisy, lagged data. ERA was tasked with optimizing the Weighted Interval Score (WIS).
- Performance: ERA's "Google Retrospective" model achieved an average WIS of 26, significantly beating the CDC’s CovidHub Ensemble (WIS 29), which is literally a robust aggregate of dozens of expert teams.
- Why it won: It moved beyond simple autoregression by hybridizing epidemiological theory with machine learning, discovering a baseline that balanced seasonal trends with short-term deviations.
Discussion: Why Tree Search Matters
The paper provides a critical insight: ERA vs. Best-of-N. Simply prompting an LLM 1,000 times and taking the best result (Best-of-N) is far less effective than using Tree Search. TS allows for meaningful iteration; the model learns from the errors of previous nodes, whereas Best-of-N is just independent sampling.
| Model | Method | Batch integration (higher better) | Epidemiology (lower better) |
|---|---|---|---|
| Gemini 3.1 Pro | Best-of-N | 0.6461 | 92.39 |
| Gemini 3.1 Pro | ERA (TS) | 0.6641 | 72.70 |
Table: Data confirms that regardless of the underlying LLM, Tree Search (ERA) provides a significant performance delta over raw sampling.
Conclusion: The Precipice of Acceleration
ERA isn't just about writing code; it's about autonomous empirical optimization. While the authors admit limitations—ERA excels at scorable tasks rather than building new physical theories from scratch—the implications for areas like drug discovery, geospatial monitoring, and numerical physics are profound.
By reducing the time to test research ideas from months to hours, ERA represents a fundamental shift in the scientific workflow: from "how do we implement this?" to "how do we score this?".
Key Takeaway: The future of scientific discovery lies in our ability to define scorable metrics for reality and then letting AI-driven tree searches tirelessly iterate on the solution.
