Agents’ Last Exam: Why Your AI Isn't Impacting the GDP Yet

Agents' Last Exam

2026-01-01
Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wanzhe Liao, Chengzhi Liu, Junbo Peng, Haoran Sun, Zechen Xu, Bo Chen, Jiayi Cheng, Yi Jiang, Keying Kuang, Yuan Li, Youbang Pan, Ziyan Rao, Alexander Schubert, Yifan Shen, Vincent Siu, Xiatao Sun, Kangqi Zhang, Xiaopan Zhang, Yuchen Zhu, Ishaan Singh Chandok, Lei Ding, Jingxuan Fan, Andrew Glover, Jiaming Hu, Yiran Hu, Wenbo Huang, Zixin Jiang, Haoran Jin, Lukas Kim, Ming Liu, Yang Liu, Alireza Rafiei, Xuhuan Shen, Kunyang Sun, Sophia Sun, Ting Sun, Eric Wang, Yixin Wang, Hanwen Xing, Sihan Xu, Yuzheng Xu, Zhongxing Xu, Zhiling Yan, Boqin Yuan, Ruiqi Zhang, Yifan Zhang, Zibo Zhao, Liana, Santanu Bosu Antu, Haoyue Bai, Carlo Bosio, Joseph Cavanagh, Patricia Cavazos-Rehg, Tianxing Chen, Xuewen Chen, Yipu Chen, Zhu Chenyu, Chen Dai, Stefano De Castro, Yunfu Deng, Kaustubh Dhole, Jiayuan Ding, Chenchen Du, Zhehang Du, Hao Fan, Run-ze Fan, Hengyu Fu, Shi Gu, Yifan Gu, Charlie Guo, Baihe Huang, Baixiang Huang, Rimika Jaiswal, Zhihan Jiang, Ran Jin, Erin Kasson, Xin Lan, Joseph Lee, Deren Lei, Chenyu Li, Daofeng Li, Haitao Li, Hongwei Li, Jingyan Li, Xiao Li, Yi Li, Yinsheng Li, Yuangang Li, Zhixu Li, Wenyu Liang, Longtai Liao, Kevin Qinghong Lin, AndyZeyi Liu, Che Liu, Jiaming Liu, Kaiyuan Liu, Xuan Liu, Pan Lu, Wenbo Lv, Yicheng Lv, Qiuyang Mang, Kyle Montgomery, Yuzhou Nie, Ruoxi Ning, Jorin Overwiening, Xu Pan, Layna Paraboschi, Core Francisco Park, Justin Purnomo, Swati Rajwal, Scott Rankin, Bixuan Ren, Yiren Rong, HaoYang Shang, Ventus Shaw, Fiona Shen, Jiawei Shen, Minqi Shi, Qiu Shi, Huaxiu Yao, Tianneng Shi, Jonah So, Vladislav Susoy, Hannah Szlyk, Haocheng Wang, Jialu Wang, Wei Wang, Xinyu Wang, Zehao Wang, Dowling Wong, Angela Wu, Dehao Wu, Fangyu Wu, Mengyuan "Millie" Wu, Yu Wu, Yuchen Wu, Yuhao Wu, Qingpo Wuwu, Weihang Xiao, Yongyi Xiong, Fan Xu, Ruiling Xu, Mingxuan Yan, Benjamin Yang, Jirong Yang, Sen Yang, Xiaoli Yang, Yushi Yang, Haoran Ye, Xiaohu Yu, Zhengming Yu, Chenlong Zhang, Chi Zhang, Hanning Zhang, Hanwen Zhang, Junge Zhang, Kunpeng Zhang, Song Zhang, Wenjin Zhang, Wenshuo Zhang, Ying Zhang, Yizhi Zhang, Brian Zhao, Qijian Zhao, Yimin Zhao, Yuhaohua Zheng, Liwei Zhou, Tianyue Zhou, Sichen Zhu, Siqi Zhu, Yan Zhu, Yishu Zhu, Jierui Zuo, Chonghao Cai, Helena Casademunt, Wenjia Chen, Benjamin Cheng, Nawen Deng, Rao Fu, Tianfu Fu, Yifan Han, Ren He, Zhenyu He, Qiao Jin, Lang Lang, Yuetai Li, Sylvia Liu, Lu Lu, Qing Lu, Subhabrata Mukherjee, Yunqi Ouyang, Yin Ren, Dawei Shi, Haoran Wu, Zhiyue Wu, Hannah Yao, Zhuoran Yi, Jenny Yu, Rhea Zhan, Hang Zhou, Blake Zhu, Junfan Zhu, Alan Yuille, Yang Liu, Russell Alan Poldrack, Jiachen Li, Zhenglu Li, Molei Tao, Jing Huang, Wenqi Shi, Costas Spanos, Lichao Sun, Chenguang Wang, Orson Xu, Zhen Dong, Hector Gomez, Aylin Caliskan, Ali Emami, Haimin Hu, Zhi Li, Lihui Liu, Murphy Niu, Yi Shao, Jianxin Sun, Mikko Tolonen, Ting Wang, Sanjiv Das, Yanjun Gao, Wenbo Guo, Erika J Schneider, Zhiyong Lu, Mark Mueller, Radha Poovendran, Somayeh Sojoudi, Dawn Song
Summary
Problem
Method
Results
Takeaways
Abstract

Agents' Last Exam (ALE) is a groundbreaking benchmark featuring 1,490 task instances across 55 subfields of digital work, designed to evaluate Generalist Computer-Use Agents (GCUAs) on long-horizon, economically valuable workflows. Using a "GUI-as-Tool" framework, it establishes a new SOTA evaluation where even frontier models like GPT-5.5 underperform, yielding an average full pass rate of only 2.6% on the hardest tasks.

TL;DR

While AI models are breaking records in math and coding, they are largely failing to perform the complex, multi-day workflows that drive the global economy. Agents’ Last Exam (ALE) is a massive new benchmark from UC Berkeley and partners that tests AI on 1,000+ real-world professional tasks. The result? Even the most advanced models (like GPT-5.5) fail over 90% of the hardest professional exams, proving that we are still far from "autonomous employees."

The Utility Gap: Benchmarks vs. Reality

We are currently witnessing a paradox: AI performance on benchmarks like MMLU or HumanEval is nearing saturation, yet most companies find that deploying these models for end-to-end professional work—like financial auditing or industrial manufacturing—remains elusive.

The authors of ALE argue this is an evaluation problem. Most benchmarks test "actions" (clicks, single lines of code), not "workflows" (completing a deliverable over hours or days). To bridge this, ALE crowdsourced projects from 250+ industry experts, covering everything from radiological adjudication to G-code generation for 3D machining.

Methodology: The Generalist Computer-Use Agent (GCUA)

ALE moves beyond the "chatbot" paradigm by defining the Generalist CUA. Unlike CLI-only agents (which can't see) or GUI-only agents (which can't code), a GCUA integrates:

  • Brain: Reasoning and planning.
  • Eyes: GUI perception via screenshots.
  • Hands: Tool invocation (files, APIs, terminal).
  • Feet: The runtime environment (Virtual Machines).

ALE Pipeline Architecture

The benchmark uses a Decoupled Architecture, separating the task specification from the agent harness. This allows for rigorous, reproducible testing where an agent is dropped into a "dirty" environment (Real Windows/Linux VMs with professional software installed) and must produce a verifiable artifact.

The "Last Exam" Frontier: Experimental Results

The benchmark is divided into three tiers: Near-Term, Full-Spectrum, and Last-Exam. The results are a wake-up call for the industry.

Performance Table

Key findings include:

  1. Low Pass Rates: Even the best model/harness combinations (GPT-5.5 + Codex) only manage a 26.2% overall pass rate.
  2. Domain Sensitivity: Models perform relatively well in "code-adjacent" fields like Computational Math (~60% score) but struggle immensely in specialized visual fields like Visual Media and Education (<30%).
  3. Knowledge vs. Execution: Interestingly, 78% of agent failures weren't due to clicking the wrong button. Instead, they were Understanding failures—the agent simply didn't know the domain-specific logic required to finish the task.

Why Conventional Agents Fail

The failure taxonomy in ALE reveals a critical "Inductive Bias" problem. When faced with a task requiring professional software (like DaVinci Resolve or SolidWorks), agents often default to writing ad-hoc Python scripts rather than using the specialized tools intended for the job. This "tool under-utilization" is a primary reason for the low scores on the harder tiers.

Failure Taxonomy

Conclusion: The Road to GDP-Relevant AI

ALE isn't just another leaderboard; it's a roadmap. It shifts the goalposts from "Abstract Intellect" to "Applied Professionalism." For AI to truly transform the economy, we must move beyond general reasoning and start training agents that can navigate complex software ecosystems with the same nuance as a human expert.

The Takeaway: If a model can pass the "Last Exam," it's no longer just a chatbot—it’s a workforce.


Note: ALE is designed as a "living benchmark," meaning the private task pool will continuously rotate to prevent data contamination from LLM pre-training.

Find Similar Papers

Try Our Examples

  • Search for recent papers that focus on "long-horizon" agent evaluation in professional or industrial software environments beyond standard web navigation.
  • Which paper first established the Standard Occupational Classification (SOC) system as a basis for AI benchmarks, and how does ALE's methodology improve upon those foundations?
  • Explore research that applies Generalist Computer-Use Agent (GCUA) architectures to specialized scientific fields like computational chemistry or structural biology.
Contents
Agents’ Last Exam: Why Your AI Isn't Impacting the GDP Yet
1. TL;DR
2. The Utility Gap: Benchmarks vs. Reality
3. Methodology: The Generalist Computer-Use Agent (GCUA)
4. The "Last Exam" Frontier: Experimental Results
5. Why Conventional Agents Fail
6. Conclusion: The Road to GDP-Relevant AI