LLMs Can't Jump: Why the Abductive Leap is the Final Frontier of AI Discovery
LLMs can't jump
This position paper by Tom Zahavy (Google DeepMind) argues that Large Language Models (LLMs) are structurally unable to achieve true scientific discovery, using Einstein's General Relativity as a case study. It distinguishes between Induction/Deduction, which AI has mastered, and Abduction, the creative "Jump" from sensory experience to new axioms, which remains an AI bottleneck.
TL;DR
In a provocative new position paper, Google DeepMind’s Tom Zahavy argues that despite their mastery of language and logic, Large Language Models (LLMs) lack the "Abductive Jump" necessary for true scientific invention. While AI can compress data (Induction) and prove theorems (Deduction), it cannot yet recreate the intuitive leap Albert Einstein took to formulate the axioms of General Relativity—a process rooted in embodied physical simulation rather than symbolic manipulation.
The Crisis of the "Chinese Room" in Science
Modern AI progress is often framed as a triumph of Inductive Inference. Researchers like Schmidhuber suggest that discovery is simply data compression—finding a simpler program to explain observations. Meanwhile, systems like AlphaProof have shown that AI is rapidly conquering Deductive Inference, the formal derivation of truths from established premises.
However, Zahavy argues that scientific invention is not just the sum of these two. Using Einstein’s journey as a case study, the paper highlights a glaring void. Einstein didn't have a "big data" problem; Newtonian gravity was incredibly accurate (within precision). There was no massive error signal driving a gradient descent. Instead, Einstein performed a mental "Jump" from sensory perception to a new system of axioms.
Methodology: The Three Pillars of Inference
The paper adopts the framework of Charles Sanders Peirce to categorize AI's missing link:
- Deduction (Rule + Case → Result): Applying a known law to a specific situation. (AI status: SOTA - Achieved)
- Induction (Case + Result → Rule): Spotting patterns in data to find a general rule. (AI status: SOTA - Achieved)
- Abduction (Rule + Result → Case/New Rule): Inventing a hypothesis to explain a surprising or singular phenomenon. (AI status: The Missing Jump)
The Bottleneck: The "Jump" (J)
Einstein’s cycle of discovery involves moving from Sense Experience (E) to a System of Axioms (A) via a Jump (J).

The paper argues that LLMs are currently "Chinese Rooms"—they manipulate the language of physics (symbols) without any access to the physical referents (sensory reality) that give those symbols meaning.
Why Data Compression and Logic Aren't Enough
Zahavy provides two critical rebuttals to current AI trends:
- Scarcity of Data: In 1907, there was no supervised training set for General Relativity. An inductive AI would have simply "patched" Newton's equations (e.g., the Vulcan planet hypothesis) rather than reinventing the geometry of spacetime.
- The Intent of Deduction: While a model like GPT-5 might derive Mercury's orbit if given Einstein's field equations, it cannot generate the Equivalence Principle (the idea that gravity and acceleration are indistinguishable) because that principle came from an embodied thought experiment, not a logical syllogism.
The Solution: Embodied World Models
How do we teach an AI to "jump"? The paper points toward Manipulative Abduction. Einstein's "Happiest Thought" came from simulating the feeling of being in a falling elevator.

The author proposes that for AI to reach this level, it requires:
- Interactive World Models: Systems like Genie that allow for action-controllability.
- Latent Physics Manifolds: Reasoning within a space of consistent physical laws rather than just predicting pixels or tokens.
- Counterfactual Intervention: The ability to "cut the cable" in a simulation to see what happens, grounded in a physical prior.
Critical Analysis & Future Outlook
The core takeaway is profound: AI invention requires perception, not just reading.
Limitations: This theory primarily addresses the physical sciences. In purely abstract fields like mathematics, the "sensory grounding" might involve high-dimensional topology or formal landscapes rather than physical gravity.
Future Work: The paper sets a new North Star for AI research. We must move beyond "AI Scientists" that recombine existing symbolic concepts and toward agents that can "experience" a synthetic laboratory. Only then can we mechanize the feedback loop between intuition and formal logic.
Final Takeaway
If we want an AI that can discover a new General Relativity, we don't just need a bigger LLM; we need an AI that can "feel" the floor of a falling elevator.
