PaSa: Beyond Keywords — Building an Autonomous Agent for Deep Academic Discovery
Pasa: An llm agent for comprehensive academic paper search
PaSa is an autonomous LLM-based academic paper search agent that utilizes two specialized components: a Crawler for multi-step citation network navigation and a Selector for precise relevance filtering. Built on the AGILE reinforcement learning framework, PaSa-7B achieves SOTA performance, significantly outperforming GPT-4o and specialized search engines in complex scholarly retrieval tasks.
TL;DR
Conducting a comprehensive literature review is one of the most time-consuming tasks for researchers. While tools like Google Scholar are ubiquitous, they struggle with "long-tail" specialized queries. PaSa (Paper Search agent) changes this by mimicking human research behavior: it doesn't just search; it reads, follows citation trails, and filters results. Built on a 7B model and optimized via Reinforcement Learning, it outperforms GPT-4o and Google Scholar by massive margins—up to 40% higher recall in real-world scenarios.
The Problem: The "Shallow Search" Trap
Why is scholarly search hard for current AI?
- Complexity: Queries like "UCB-based algorithms for non-stationary RL in value-based methods" are too granular for simple keyword matching.
- Missing Links: A single search query often misses seminal works that are only discoverable by reading a paper's citations (the "citation network").
- Information Overload: LLMs often hallucinate or provide surface-level relevance, failing to verify if a paper truly addresses the specific constraints of a query.
Methodology: Crawler meets Selector
PaSa operates via two distinct LLM agents, a structure that mirrors the human research process of discovery and verification.
1. The Crawler (Discovery-Centric)
The Crawler is the "explorer." It manages a paper queue and uses three primary actions:
- [Search]: Generates diverse queries to find the initial batch of papers.
- [Expand]: Deep-dives into specific sections of a paper to identify relevant citations.
- [Stop]: Signals the completion of a session to move to the next paper.
2. The Selector (Precision-Centric)
The Selector is the "judge." It performs deep reading of abstracts and rationale generation. Crucially, during the training phase, the Selector serves as an Auxiliary Reward Model. This solves the "sparse reward" problem—even if a paper found by the Crawler isn't in the ground-truth dataset, the Selector can reward the Crawler for finding something relevant.

Reinforcement Learning: Teaching an Agent to "Browse"
The authors utilized the AGILE framework and introduced Session-level PPO. Because a full search trajectory can include hundreds of papers (far exceeding context windows), they broke the task into manageable sessions. By applying a reward function that penalizes excessive actions (action cost) but rewards finding target papers, they successfully taught a 7B model to navigate efficiently.
Experimental Results: Dominating the Baselines
The researchers developed two benchmarks: AutoScholarQuery (35k synthetic pairs from top AI conferences) and RealScholarQuery (annotated by professors from top-tier universities).
| Method | Recall@20 (Real) | Precision (Real) |
|---|---|---|
| Google Scholar | 15.14% | - |
| Google + GPT-4o | 20.20% | - |
| GPT-o1 (No search) | 1.34% | 5.80% |
| PaSa-GPT-4o (Prompting) | - | 47.21% |
| PaSa-7b (RL Optimized) | 57.98% | 51.46% |
The results show a stark reality: even highly capable models like GPT-o1 fail at academic search without autonomous tool use and citation crawling. PaSa-7b's ability to achieve 61.11% overall recall compared to GPT-4o's 30.75% emphasizes that RL training is far more effective than just writing better prompts.

Critical Insights & Takeaways
- Inductive Bias of the Citation Network: The most powerful feature of PaSa is the
[Expand]action. The ablation study shows that removing citation expansion drops recall by over 32%. Academic knowledge is inherently a graph; searching it linearly is suboptimal. - Small Models, Big Performance: The fact that a 7B model (Qwen2.5) can outperform GPT-4o highlights that for specialized tasks, fine-tuning/RL on domain-specific workflows is more important than raw parameter count.
- Limitations: Currently, PaSa is optimized for the AI/ML field. The cost of running multi-step agents remains higher than a single search, though the time saved for a human researcher likely outweighs the compute cost.
Conclusion
PaSa shifts the paradigm of academic search from "Retrieval" to "Research." By providing an agent that can navigate citation networks and explain its reasoning, this work lays the groundwork for fully automated scientific literature surveys.
For those interested in trying the tool or viewing the code, visit the PaSa GitHub repository.
