Theory Is All You Need? The Data-Belief Asymmetry Between LLM Prediction and Human Causal Reasoning
Theory is all you need: AI, human cognition, and causal reasoning
Teppo Felin and Matthias Holweg (2024) argue that large language model prediction and human cognition are fundamentally different: LLMs extrapolate from past data through statistical association, while humans can form forward-looking theories that drive causal reasoning, experimentation, and new evidence. The paper introduces data-belief asymmetry as the condition under which beliefs outpace existing data, thereby enabling novelty and new knowledge. Its central conclusion is that AI is not a general substitute for human decision making under uncertainty.
TL;DR
Teppo Felin and Matthias Holweg (2024) argue that artificial intelligence and human cognition are not instances of the same computational process. Large language models, they contend, are extraordinarily capable prediction engines whose outputs are drawn from statistical associations in past text, while human cognition is fundamentally forward-looking because it can form theories that are not fully justified by current data. The paper introduces the notion of data-belief asymmetry to explain how beliefs can outpace evidence and become engines of novelty, intervention, and new knowledge. Its main conclusion is that AI will remain powerful for routine, data-rich prediction, but it cannot by itself replace human judgment in genuinely uncertain strategic settings.
Positioning in the Literature
This is a position paper in strategy, published in Strategy Science, volume 9, issue 4, pages 346-371. It belongs to a contrarian theoretical genre that challenges the equivalence between computation and cognition. The authors do not propose an algorithm, benchmark, or experimental model. Their contribution is conceptual: they argue that the long-standing analogy between minds and machines as input-output information processing devices breaks down exactly where novelty, counterfactual reasoning, and directed experimentation enter the picture.
The paper is situated against two broad claims. One claim is empirical and comes from the AI performance narrative: large language models pass professional exams, diagnose, plan, negotiate, and sometimes outperform humans. The other claim is more theoretical and comes from computational cognitive science: cognition is a form of information processing, prediction, or probabilistic inference. Felin and Holweg do not deny the first claim. They accept that LLMs are powerful and that some tasks can be automated. What they reject is the inference that these capabilities reveal the same underlying cognitive mechanism that humans use when they generate new knowledge and act under uncertainty.
The work is also a descendant of the theory-based view in strategy, which the authors themselves have developed in prior work. The paper's title echoes the famous transformer paper, but the authors explain that they do not mean theory is literally all humans need. They mean that theory is a foundational and often unrecognized dimension of cognition, especially in situations where evidence is incomplete, misleading, or yet to be created.

Problem and Motivation
The paper's starting point is that AI progress has intensified an old ambition: to model thinking itself as computation. From the Dartmouth conference onward, researchers have treated intelligence as describable in terms of inputs, representations, algorithms, and outputs. The paper also notes that this idea has spread far beyond early symbolic AI. It appears in neural networks, Bayesian models of cognition, predictive processing, active inference, and computational theories of mind. In all of these, the shared premise is that cognition is information processing and that learning is the internalization of structure from data.
The authors identify a precise failure mechanism: this framework cannot easily explain where genuinely new knowledge comes from. If belief should be proportional to evidence, and evidence is the product of past observations, then the cognitive system is structurally backward-looking. It can represent, recombine, summarize, and translate what is already in the environment. But it has no intrinsic mechanism for projecting into a future that is not already statistically encoded in the data.
This problem becomes especially severe in uncertain settings. Routine prediction is useful when the future resembles the past. But strategic decisions often involve rare, high-impact possibilities that are not represented in current data. The paper argues that prediction cannot solve this problem because it starts from the given world, whereas novelty requires asking about a not-yet-existing world. The relevant question is not only what the current evidence supports, but what actions or experiments would make new evidence observable.
The paper therefore reframes the debate. It does not merely say AI is imperfect. It says that data-based prediction and theory-based causal reasoning are different kinds of epistemic activity. The first updates beliefs from evidence. The second can create the conditions under which evidence appears.
The Argument: From Prediction to Causal Intervention
The Input-Output Premise: Why LLM Fluency Is Backward-Looking
The paper's first core move is to reconstruct how a large language model learns and to contrast that process with human language acquisition. LLMs are trained on enormous corpora of written text. The authors explain that modern models are estimated to incorporate roughly 13 trillion tokens, where a token is roughly equivalent to a word. They also note that if a human read at 9,000 words per hour, the same corpus would require more than 164,000 years. The point is not simply that machines have more data. The point is that their learning architecture is designed to discover statistical associations among those data.
In the paper's description, language is tokenized into numerical representations. Each token is embedded into a dense vector that captures semantic and positional information. The transformer architecture, introduced in the paper as the breakthrough behind modern LLMs, then allows tokens to be processed in relation to other tokens. The final result is a model that generates text through stochastic next-word prediction. It samples from conditional probabilities derived from patterns in the training corpus.
This description supports the authors' broader argument: the model's intelligence is not absent, but it is specialized. LLMs are fluent because the same content can be expressed in many ways. They are creative at the level of recombination because they can produce novel sentences by sampling from combinatorial possibilities already latent in the training distribution. The paper describes this as a lowercase generativity: creative summarization, rephrasing, and repackaging rather than genuine knowledge generation.
The authors use the example of language learning to challenge the symmetry between human and machine cognition. LLMs are trained on trillions of written tokens. Children are exposed to far less language, often spoken, often messy, and often incomplete. The paper notes that an infant hears about 20,000 words per day, which amounts to roughly 36.5 million words over five years. By contrast, if a child were exposed to 13 trillion tokens, it would take roughly 1.8 million years, according to the paper's footnote. The human linguistic output is therefore underdetermined by the input. A child does not simply memorize and recombine heard sentences; it generates grammatical and interpretive structures that are not point-by-point contained in experience.
| Reported dimension | Figure in the paper | Function in the argument |
|---|---|---|
| LLM training corpus | roughly 13 trillion tokens | Shows the scale of machine input |
| Human reading time for that scale | more than 164,000 years at 9,000 words per hour | Makes machine-scale data tangible |
| Infant spoken exposure | about 20,000 words per day, or about 36.5 million words over five years | Contrasts human data scarcity |
| Infant exposure to match LLM scale | roughly 1.8 million years, per footnote | Reinforces underdetermination of human language |
| Exam performance | more than 90 percent of humans in some professional exams | Sets up the tempting inference of cognitive equivalence |
The authors use these quantities not as experimental results, but as conceptual contrasts. The table helps isolate the paper's main intuition: high performance can arise from statistical exposure and recombination, but it does not automatically imply the same mechanism that lets a child construct grammar from impoverished input or lets a scientist construct a theory from incomplete evidence.
The paper extends the analogy beyond language. It argues that LLMs share with predictive-processing approaches a conservative logic of error minimization. Both try to reduce the mismatch between prediction and outcome. An LLM minimizes prediction error over words; the predictive brain minimizes surprise about perception. In the paper's framing, this shared logic is precisely what limits novelty. A system designed to minimize error from existing data is not naturally oriented toward generating new data that contradicts the existing pattern.
This is the first place where the authors' design reasoning matters. Why choose LLMs rather than some other AI system? The paper argues that LLMs are the current exemplar of data-driven prediction, and that many AI-related claims about novelty and reasoning are made specifically about language models. By taking the most visible case, the authors can attack the broader premise that scale plus statistics equals cognition.
Data-Belief Asymmetry: The Missing Mechanism for Novelty
The central conceptual contribution of the paper is the distinction between data-belief symmetry and data-belief asymmetry. Data-belief symmetry is the epistemic stance according to which belief should be proportionate to evidence. This is the stance of Bayesian updating, probabilistic cognition, and many machine learning formulations. In the paper's description, knowledge is justified belief, and belief is justified by data and evidence. The stronger the evidence, the stronger the belief.
The authors acknowledge that this assumption explains a large amount of rational cognition. It also explains a large amount of error. The cognitive sciences have heavily studied the negative side of asymmetry: humans persist in beliefs despite contrary evidence, selectively sample data, display confirmation bias, motivated reasoning, availability bias, and many other pathologies. The paper notes that this downside is well represented in the literature and has influenced economics, judgment research, and behavioral strategy.
The paper's innovation is to flip the evaluation. What if some data-belief asymmetries are not cognitive failures but preconditions of knowledge creation? A belief that outstrips current evidence may look delusional ex ante. Yet if it motivates targeted intervention and experimentation, it may produce evidence that would never have appeared otherwise. This is the positive side of asymmetry.
The authors connect this to the problem of relevance. Data alone does not tell a decision maker which data matter. They quote the paper's own formulation that things are not labeled evidence in nature. A rational system that only updates from available evidence must still be told what counts as relevant evidence. But in genuinely novel domains, the relevant evidence may not yet exist. It must be generated by an intervention guided by a theory.
The Galileo thought experiment dramatizes this point. Imagine an LLM in 1633 trained on all human text up to that time. Asked about heliocentrism, it would likely mirror the dominant geocentric consensus. The overwhelming frequency of geocentric associations in the training data would not merely reflect past belief; under a statistical epistemology, it would look like evidence. The paper argues that LLMs cannot distinguish Galileo's then-contrarian belief from Tycho Brahe's astrological claims on the basis of truth, because they have no mechanism for access to reality beyond textual frequency.
This thought experiment is more vivid than persuasive for a modern AI researcher. The paper could have answered a stronger objection: large models also include minority reports, anomalies, and conflicting signals. But the authors' point is structural. A purely statistical system has no ex ante method for deciding which minority signal is prescient and which is false. It can reweight frequency, retrieve expert sources, or combine models, but these are still ways of summarizing what has already been written. The paper discusses mixture-of-experts models, retrieval-augmented generation, and ensemble methods as partial attempts to improve reliability. Its objection is that these methods still condition on human-authored corpora and selected sources. They can improve accuracy, calibration, or domain coverage, but they do not automatically give the model forward-looking causal logic.
The Wright brothers example provides the paper's clearest illustration of the asymmetry. In the late nineteenth and early twentieth centuries, much of the available evidence and expert consensus argued against heavier-than-air human flight. Joseph LeConte, described by the paper as a prominent scientist who later became president of the American Association for the Advancement of Science, argued from bird data that flight was impossible above a certain weight. The paper notes his observation that no bird above 50 pounds flies. Simon Newcomb, described as a foremost astronomer and mathematician, argued that the largest flying birds rise with difficulty and that even the condor, lighter than a human, struggles when gorged with food. Lord Kelvin stated that heavier-than-air flying machines were impossible. The New York Times estimated in 1903, after Samuel Langley's failure, that human-powered flight might take from one million to ten million years of combined effort.
The point is not that these scientists were foolish. They were doing exactly what a data-belief-symmetric system does: forming beliefs from available evidence and weighting that evidence by source reliability. The paper argues that the Wright brothers were successful because they did not stop at that logic. Their belief in flight was not a conclusion derived from existing data. It was a forward-looking commitment that motivated causal reasoning. They asked what had to be true for flight to be possible, and then designed experiments to make that truth discoverable.
This is the most essential differentiator in the paper: novelty is not explained by updating belief from data. It is explained by using belief to generate a theory, using theory to identify interventions, and using interventions to create evidence. The authors emphasize that this is not a data-independent fantasy. It is a feedback process in which human agents decide which experiments to run, what to measure, and how to interpret failure.
Causal Reasoning and Experimentation: Where Beliefs Create Evidence
The third core section translates the historical example into a cognitive mechanism. The authors define theory-based cognition as a forward-looking activity. A theory does not merely mirror reality; it proposes a path from the current state of the world to a hypothesized future state. The paper ties this to Pearl's interventionist causality: causal reasoning is not only about predicting associations, but about asking what would happen if the system were manipulated.
This distinction matters because prediction and intervention are different cognitive tasks. A prediction system asks what the next likely outcome is given the data. A causal intervention system asks what must be changed, what must be measured, and what failure patterns reveal necessary and sufficient conditions. The paper argues that LLMs are built for the first task and lack the second.
The Wright brothers' work on lift illustrates the mechanism well. The paper explains that they studied historical flight attempts, including Otto Lilienthal's records. They did not simply aggregate prior data. They interrogated it for causal structure. They developed experiments with airfoils, tested wing shapes, sizes, and angles, and constructed their own wind tunnels. This allowed them to generate new aerodynamic data under controlled conditions. Their eventual discovery and use of wing warping for roll control is presented as the outcome of understanding the causal relationship between wing shape, air pressure, and lift.
The paper also describes their problem formulation as decisive. They reasoned that three problems had to be solved: lift, propulsion, and steering. This is where the authors connect the argument to strategy. The capacity to formulate the right problems is not just another form of reasoning. It is a theory-driven act that creates salience for future data. Without a causal theory of flight, the relevant variables would remain invisible or mislabeled. With it, failure becomes information, and experimentation becomes a method for discovering what must be true for the belief to be realized.
The paper extends this to the choice of Kitty Hawk. Before attempting flight, the Wright brothers consulted the U.S. Weather Bureau because they had a theory of the conditions needed for controlled testing: consistent wind, open space, soft landing surfaces, and privacy. This detail supports a broader claim: forward-looking decisions are often less about extracting predictions from existing data than about arranging the environment so that decisive evidence can be produced.
The authors then turn this into an argument against prediction-centered decision making. They discuss the chain often associated with AI economics: data leads to information, information leads to prediction, prediction leads to decision. They argue that this chain is powerful for routine decisions, where the past is a reliable guide. But it breaks down for high-uncertainty strategy because the most valuable options are not already represented in past data. In such cases, the question is not what the data predict, but what intervention will generate data capable of testing a theory.
Here the paper reaches its strongest strategic conclusion. In competitive settings, if all actors use the same prediction engines trained on the same public data, outputs tend toward generic expectations. Value creation requires firm-specific theories, proprietary data, and unique causal paths. The authors suggest that retrieval-augmented generation and customized fine-tuning can help make AI more specific, but only if humans deliberately decide which corpora matter, which sources to include or exclude, and which hypotheses the system is being asked to support.

Evidence, Experiments, and Limits
This paper does not report an experiment, benchmark, or ablation study. That is an important qualification. Its evidence is conceptual and historical. The quantitative figures are illustrative, not experimental. The table above shows the paper's main numerical anchors, but those numbers support a thought experiment about input scale and underdetermination, not a controlled comparison of reasoning mechanisms.
The strongest empirical claims in the paper come from cited AI performance findings, such as the statement that current AI models outperform more than 90 percent of humans in some professional exams. The authors use these facts to show why the computational analogy is tempting. But they also use them to sharpen the distinction between performance and mechanism. A model can pass an exam without possessing the kind of causal logic that would let it decide, before the fact, which minority hypothesis will later become knowledge.
The historical evidence is also selective. The paper openly states that it has opportunistically selected a case in which a belief contrary to existing evidence turned out to be correct. This honesty is a virtue, but it also limits inference. The Wright brothers case illustrates a mechanism, but it does not provide a statistical proof that most delusional beliefs become successful innovations. In fact, the paper's argument is that the mechanism is conditional on causal reasoning and experimentation. The Wright brothers succeeded because they built a theory of flight and tested it. Not every agent who ignores evidence does so.
The conceptual evidence is stronger than the historical case. The paper's account of LLMs makes sense within its own framework: next-word prediction produces fluent recombination of past text, and therefore struggles to explain the origin of novel truth. Its account of human theory-based cognition is also internally coherent: beliefs can motivate action, action can generate new evidence, and evidence can later justify what was initially unjustified by existing data.
The paper's main weakness is that it leaves open the boundary between productive and pathological asymmetry. It says that beliefs can be measured by propensity to act, but it does not formalize when action is a causal intervention versus confirmation bias. The Wright brothers had a theory that guided measurable experiments. A conspiracy theorist may also act on belief while selectively gathering evidence. The paper distinguishes these cases by invoking causal reasoning and experimentation, but it does not provide a criterion that can be applied without judgment. That absence matters because strategy is exactly the domain where judgment is scarce.
A second limitation concerns the speed of AI change. The paper argues that current LLMs are backward-looking, and for present architectures that is fair. But it also acknowledges, in its endnotes, that future systems might mimic forms of reasoning the authors currently treat as uniquely human. The Galileo thought experiment assumes an LLM whose entire epistemic access is textual frequency. Modern research may add tools, retrieval, simulation, planning, or feedback. The paper's argument remains valid for those components insofar as they are still selected and framed by human theories, but it should not be read as a claim that future AI can never participate in discovery.
A third limitation is the paper's relative silence on formal causal AI. It contrasts LLMs with human causal reasoning and invokes interventionist causality, but it does not engage deeply with structural causal models, causal graphs, experimental design algorithms, or causal discovery systems. That omission may be deliberate because the target is the computational view of cognition and LLM hype. Still, the argument would be stronger if it separated more carefully between statistical prediction and formal causal modeling. A causal model is not merely a language model; it can represent interventions and counterfactuals even if it is implemented computationally.
Insights and Conclusion
The paper's deepest insight is that novelty requires a cognitive structure that is not reducible to evidence updating. Data-belief symmetry can explain calibration, but not the origin of new relevance. If evidence is needed to justify a belief, but the belief is required to discover the evidence, then there must be a different mode of cognition. The authors name that mode theory-based causal reasoning. It is a process in which a forward-looking belief identifies problems, guides counterfactual reasoning, and structures experimentation.
The contribution to strategy is specific. AI should not be understood as a replacement for strategic judgment in high-uncertainty domains. It is better understood as a tool for compressing, summarizing, and recombining existing knowledge. If firms use AI generically, they risk producing the same generic outputs as competitors. If they use it purposefully, they may use it as a custom instrument embedded in a firm-specific theory of value. The paper suggests retrieval-augmented generation and proprietary data as possible paths, but it also insists that the selection of data remains a human theoretical act.
The conclusion follows from that. LLMs are impressive because they turn past language into fluent present output. Human knowledge advances because agents can act on beliefs that are not yet justified by past data. The difference is not merely one of scale, noise, bias, or processing power. It is a difference between systems that update from the world and systems that intervene in the world to make new evidence. The paper's title is therefore provocative but precise. Theory is not literally all you need. But without theory, the data needed to justify a new truth may never be produced.
