文章背景与核心概要
本文探讨了大语言模型(LLM)在溯因推理(abductive reasoning)方面的根本局限性。溯因推理是实现科学突破(如爱因斯坦的等效原理或普朗克的黑体辐射定律)的核心认知能力。与以往认为“具身智能”缺失是主要障碍的观点不同,作者提出问题的根源在于认知过程缺乏“认知错误”与“物理代价”之间的耦合。
研究通过分析普朗克推导 \(E = h\nu\) 的历史案例,指出真正的科学溯因需要一种机制,使得认知错误能够产生足以迫使模型进行根本性修正的物理代价。作者通过热力学耦合理论证明,无论模型规模如何扩大,固定权重的 Transformer 推理架构都无法实现这种耦合。实证数据表明,即便任务的因果难度显著增加且准确率大幅下降,模型的输出熵依然保持不变,这进一步证实了当前 AI 架构在处理深层科学推理时的结构性缺陷。
LLMs Don't Pay for the Jump
Summary
大语言模型(LLM)在溯因推理方面存在根本性局限,即无法实现科学重大突破所需的“跨越”(例如爱因斯坦的等效原理或马克斯·普朗克对黑体辐射问题的解决)。
This paper investigates the fundamental limitations of Large Language Models (LLMs) regarding abductive reasoning—the capacity to make the conceptual "Jumps" characteristic of profound scientific breakthroughs (such as Einstein's equivalence principle or Max Planck's resolution of the blackbody radiation problem).
尽管一些研究人员认为这些局限性源于缺乏具身模拟,但作者认为问题更为深远。通过考察历史上的突破性进展(如普朗克推导 \(E = h\nu\) 的过程,该过程并不需要传感器运动接地),本文证明了真正的科学溯因依赖于认知错误(epistemic error)与物理代价(physical cost)之间的耦合。作者指出,无论规模如何,固定权重的 Transformer 推理都缺乏这种热力学耦合——实证数据支持了这一发现,即随着因果任务难度增加和准确率下降,输出熵依然保持基本静态。因此,作者得出结论:机器溯因需要一种机制,使认知错误产生足以迫使认知修正的真实物理后果。
While some researchers argue that these limitations stem from a lack of embodied simulation, the authors suggest the problem runs deeper. Examining historical breakthroughs (like Planck's derivation of \(E = h\nu\), which required no sensorimotor grounding), the paper demonstrates that true scientific abduction relies on a coupling between epistemic error and physical cost. The authors show that fixed-weight transformer inference lacks this thermodynamic coupling regardless of scale—a finding supported by empirical data showing output entropy remains largely static even as causal task difficulty increases and accuracy drops. Consequently, the authors conclude that machine abduction requires a mechanism where epistemic errors carry real physical consequences powerful enough to force cognitive revision.
Metadata
- arXiv ID: arXiv:2608.14397 [cs.AI]
- Subjects: Artificial Intelligence (
cs.AI); Computation and Language (cs.CL) - Authors: Paras Balani, Subhrakanta Panda
- Submitted: August 14, 2026
- Length: 14 pages
- License: Creative Commons Attribution 4.0 International

Abstract
Zahavy (2026) 认为,大语言模型尽管在归纳和演绎方面表现出色,但无法执行产生爱因斯坦等效原理的那种溯因“跨越”,并将此局限性归因于缺乏具身模拟。Zheng-Xin (2026) 和 Farmer (2026) 则质疑具身性对于溯因是否必要,并指出了通向广义相对论的其他路径,以及不需要传感器运动接地的溯因形式。
Zahavy (2026) argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simulation. Zheng-Xin (2026) and Farmer (2026) question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding.
马克斯·普朗克在 1900 年解决了黑体辐射问题。普朗克推导 \(E = h\nu\) 的过程并不需要具身模拟。其动力源于经典理论的一个数学结论——即有限测量量对应的预测能量为无穷大——这在物理上是无法接受的。
Max Planck resolved the blackbody radiation problem in 1900. Planck's move to \(E = h\nu\) required no embodied simulation. It was motivated by a mathematical consequence of classical theory—an infinite predicted energy for a finite measured quantity—that could not be physically accepted.
作者指出,归纳法和演绎法都无法产生这一假设,并认为该假设的采纳需要认知错误与物理代价之间的耦合。他们通过“热力学耦合”将这种区别形式化,并证明了无论模型规模如何,固定权重的 Transformer 推理都缺乏这种耦合。这与实证结果一致:在因果难度急剧增加的任务中,尽管准确率从 100% 下降到 17%,但输出熵几乎保持不变。
The authors show that neither induction nor deduction could have produced the postulate, arguing instead that its adoption required a coupling between epistemic error and physical cost. They formalize this distinction through thermodynamic coupling and demonstrate that fixed-weight transformer inference lacks such coupling, regardless of model scale. This aligns with empirical results showing that output entropy remains nearly unchanged across tasks with sharply increasing causal difficulty, even as accuracy falls from 100% to 17%.
最终,机器溯因缺失的要素可能比具身性更深层:系统必须具备一种物理机制,通过该机制,认知错误变得足够昂贵,从而迫使系统进行根本性的修正。
Ultimately, the missing ingredient in machine abduction may lie deeper than embodiment: a system must possess a physical mechanism through which epistemic error becomes costly enough to force radical revision.
Links & Resources
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- Citation & References:
- Google Scholar
- Semantic Scholar
- NASA ADS