What is the 'jagged frontier' and why does it make overconfidence dangerous?
The jagged frontier describes AI's uneven performance: it can excel at tasks humans find hard while failing at tasks humans find easy, even within the same workflow. This unpredictability is not random—it follows patterns that are often hard for users to learn. The concept was formalized in a large field experiment with 758 consultants at Boston Consulting Group [3], which showed that AI assistance helped workers complete 12.2% more tasks 25.1% faster with higher quality—but only for tasks within AI's capability frontier. For a complex managerial task outside that frontier, AI users were 19% less likely to produce correct solutions than those without AI. This means overconfidence—assuming AI will help when it actually hurts—directly reduces performance.
A separate organizational study [2] confirms this pattern in real-world settings, noting that AI degrades outcomes on some knowledge tasks while improving others. The danger is that users, seeing AI succeed on many tasks, may assume it is universally capable and fail to spot the tasks where it is unreliable. The jagged frontier makes this overconfidence especially risky because the failures are not obvious—they can occur on tasks that seem similar to ones AI handles well.
Do AI evaluations help users calibrate their trust?
Not reliably, according to the evidence. A preregistered experiment with 577 participants [1] tested whether showing AI errors on easy versus hard tasks changed how much users relied on the AI. The researchers expected that errors on easy tasks (which violate expectations of AI competence) would reduce use more than errors on hard tasks. Instead, they found that observing more errors reduced use overall, but easy-task errors did not reduce use significantly more than hard-task errors. This suggests that people do not naturally learn the jagged pattern—they simply become more cautious across the board, or not cautious enough in the right places.
In a medical context, a study of 14 obstetrics residents using ChatGPT [6] found that users' self-assessed AI skills did not correlate with the accuracy of AI responses they obtained. Residents showed moderate IT proficiency but low AI proficiency, and their queries produced only 21% accurate responses due to misinterpretation of medical acronyms. Despite this, the AI's responses were plausible-sounding—a manifestation of the 'stochastic parrot' phenomenon. This means even when users think they are evaluating AI output, they may be misled by its fluency. Evaluations alone, without structured AI literacy training, do not prevent overconfidence.
Do the AI models themselves know their limits?
No—frontier models are consistently overconfident about their own capabilities. A study of LLM self-knowledge [4] found that even advanced models like GPT-4o and Mistral Large are not sure of their own capabilities more than 80% of the time, meaning they lack reliable self-awareness. The models swing between overconfidence and conservatism depending on the task category, with the biggest weaknesses in temporal awareness and contextual understanding. This internal confusion means the models cannot reliably flag when they are likely to fail.
Another study [5] tested whether LLMs can predict their own success on tasks and found that all tested models are overconfident. Newer and larger models do not generally have better discriminatory power—they are not better at knowing when they will fail. On multi-step agentic tasks, overconfidence actually worsens as the models progress through steps. When given in-context experiences of failure, some models reduced overconfidence and improved decision-making, but others did not. This suggests that even if evaluations are built into the system, the models themselves may not learn from them without explicit training. The implication is that frontier AI evaluations, whether performed by humans or by the models themselves, are not a cure for overconfidence—they must be paired with structured workflows, human judgment training, and system design that accounts for the jagged frontier.
About These Sources
This answer is built on 6 peer-reviewed studies — published from 2024 to 2026, 6 from 2024 or later, 1 in Q1–Q2 journals, collectively cited 104 times — selected as the most relevant from 7 studies that passed quality screening, drawn from 36 papers retrieved from a database of over 500 million.
Sources used in this answer
Effects of Generative AI Errors on User Reliance Across Task Difficulty
In a preregistered experiment with 577 participants, observing more AI errors reduced use, but easy-task errors did not reduce use significantly more than hard-task errors, suggesting people do not naturally learn jagged error patterns.
Navigating the Jagged Technological Frontier: Organizational Strategies for AI Integration in Knowledge Work
Drawing on field experimental evidence from 758 consultants, this article reports that AI users were 19% less likely to produce correct solutions on complex tasks outside AI's capability frontier, indicating overreliance risks.
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality
In a preregistered experiment with 758 knowledge workers, AI assistance enabled 12.2% more tasks completed 25.1% faster with higher quality within the frontier, but for a complex task outside it, AI users were 19% less likely to produce correct solutions.
Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
Frontier models like GPT-4o and Mistral Large are not sure of their own capabilities more than 80% of the time, swinging between overconfidence and conservatism depending on task category, with weaknesses in temporal and contextual understanding.
Do Large Language Models Know What They Are Capable Of?
All tested LLMs are overconfident about their success on tasks; newer and larger models do not generally have better discriminatory power, and overconfidence worsens on multi-step tasks for several frontier models.
AI in obstetrics: Evaluating residents’ capabilities and interaction strategies with ChatGPT
In a study of 14 obstetrics residents using ChatGPT, only 21% of responses were accurate due to misinterpretation of medical acronyms, and there was no correlation between self-assessed AI skills and response accuracy.
