PaSa: Beyond Keywords — Building an Autonomous Agent for Deep Academic Discovery

Pasa: An llm agent for comprehensive academic paper search

2025-07-01
Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, Weinan E
Summary
Problem
Method
Results
Takeaways
Abstract

PaSa is an autonomous LLM-based academic paper search agent that utilizes two specialized components: a Crawler for multi-step citation network navigation and a Selector for precise relevance filtering. Built on the AGILE reinforcement learning framework, PaSa-7B achieves SOTA performance, significantly outperforming GPT-4o and specialized search engines in complex scholarly retrieval tasks.

TL;DR

Conducting a comprehensive literature review is one of the most time-consuming tasks for researchers. While tools like Google Scholar are ubiquitous, they struggle with "long-tail" specialized queries. PaSa (Paper Search agent) changes this by mimicking human research behavior: it doesn't just search; it reads, follows citation trails, and filters results. Built on a 7B model and optimized via Reinforcement Learning, it outperforms GPT-4o and Google Scholar by massive margins—up to 40% higher recall in real-world scenarios.

The Problem: The "Shallow Search" Trap

Why is scholarly search hard for current AI?

  1. Complexity: Queries like "UCB-based algorithms for non-stationary RL in value-based methods" are too granular for simple keyword matching.
  2. Missing Links: A single search query often misses seminal works that are only discoverable by reading a paper's citations (the "citation network").
  3. Information Overload: LLMs often hallucinate or provide surface-level relevance, failing to verify if a paper truly addresses the specific constraints of a query.

Methodology: Crawler meets Selector

PaSa operates via two distinct LLM agents, a structure that mirrors the human research process of discovery and verification.

1. The Crawler (Discovery-Centric)

The Crawler is the "explorer." It manages a paper queue and uses three primary actions:

  • [Search]: Generates diverse queries to find the initial batch of papers.
  • [Expand]: Deep-dives into specific sections of a paper to identify relevant citations.
  • [Stop]: Signals the completion of a session to move to the next paper.

2. The Selector (Precision-Centric)

The Selector is the "judge." It performs deep reading of abstracts and rationale generation. Crucially, during the training phase, the Selector serves as an Auxiliary Reward Model. This solves the "sparse reward" problem—even if a paper found by the Crawler isn't in the ground-truth dataset, the Selector can reward the Crawler for finding something relevant.

Architecture of PaSa

Reinforcement Learning: Teaching an Agent to "Browse"

The authors utilized the AGILE framework and introduced Session-level PPO. Because a full search trajectory can include hundreds of papers (far exceeding context windows), they broke the task into manageable sessions. By applying a reward function that penalizes excessive actions (action cost) but rewards finding target papers, they successfully taught a 7B model to navigate efficiently.

Experimental Results: Dominating the Baselines

The researchers developed two benchmarks: AutoScholarQuery (35k synthetic pairs from top AI conferences) and RealScholarQuery (annotated by professors from top-tier universities).

MethodRecall@20 (Real)Precision (Real)
Google Scholar15.14%-
Google + GPT-4o20.20%-
GPT-o1 (No search)1.34%5.80%
PaSa-GPT-4o (Prompting)-47.21%
PaSa-7b (RL Optimized)57.98%51.46%

The results show a stark reality: even highly capable models like GPT-o1 fail at academic search without autonomous tool use and citation crawling. PaSa-7b's ability to achieve 61.11% overall recall compared to GPT-4o's 30.75% emphasizes that RL training is far more effective than just writing better prompts.

Paper Search Example

Critical Insights & Takeaways

  • Inductive Bias of the Citation Network: The most powerful feature of PaSa is the [Expand] action. The ablation study shows that removing citation expansion drops recall by over 32%. Academic knowledge is inherently a graph; searching it linearly is suboptimal.
  • Small Models, Big Performance: The fact that a 7B model (Qwen2.5) can outperform GPT-4o highlights that for specialized tasks, fine-tuning/RL on domain-specific workflows is more important than raw parameter count.
  • Limitations: Currently, PaSa is optimized for the AI/ML field. The cost of running multi-step agents remains higher than a single search, though the time saved for a human researcher likely outweighs the compute cost.

Conclusion

PaSa shifts the paradigm of academic search from "Retrieval" to "Research." By providing an agent that can navigate citation networks and explain its reasoning, this work lays the groundwork for fully automated scientific literature surveys.


For those interested in trying the tool or viewing the code, visit the PaSa GitHub repository.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply reinforcement learning or agentic frameworks to improve specialized document retrieval beyond general web search.
  • Which paper first introduced the AGILE framework for LLM agent optimization, and how does the session-level PPO in PaSa specifically modify it for long trajectories?
  • Are there any studies exploring the application of citation-network-crawling agents like PaSa to non-AI scientific domains such as medicine or materials science?
Contents
PaSa: Beyond Keywords — Building an Autonomous Agent for Deep Academic Discovery
1. TL;DR
2. The Problem: The "Shallow Search" Trap
3. Methodology: Crawler meets Selector
3.1. 1. The Crawler (Discovery-Centric)
3.2. 2. The Selector (Precision-Centric)
4. Reinforcement Learning: Teaching an Agent to "Browse"
5. Experimental Results: Dominating the Baselines
6. Critical Insights & Takeaways
7. Conclusion