跳转至

文章背景与核心概要

尽管大型语言模型(LLM)能够准确回答科学问题,但正确的输出并不天然证明模型真正理解或表征了底层的物理机制。本文通过考察三个 Gemma 4 模型(google/gemma-4-E4B-itgoogle/gemma-4-12B-itgoogle/gemma-4-31B-it),深入研究了开源大语言模型如何表征物理学原理。

研究在这些模型中识别出三个截然不同且可通过实验分离的材料科学信息特征:1. 选择性概念可读性;2. 定性本构取向的关系编码;3. 受限工程答案的因果、上下文相关控制。通过利用匹配的直接与雅可比词汇读出(Jacobian vocabulary readouts)、无选项状态几何结构、60条定律的反事实基准测试以及因果干预,该研究证明,在受控的状态转换中,物理关系比单纯的绝对隐状态要清晰得多。


Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Authors: Markus J. Buehler
Identifiers: arXiv:2607.20058 [cs.AI] | DOI: 10.48550/arXiv.2607.20058
Submission History: Submitted on 22 Jul 2026; last revised 3 Sep 2026 (v2).

Authors: Markus J. Buehler
Identifiers: arXiv:2607.20058 [cs.AI] | DOI: 10.48550/arXiv.2607.20058
Submission History: Submitted on 22 Jul 2026; last revised 3 Sep 2026 (v2).


Executive Summary

While large language models (LLMs) can accurately answer scientific questions, a correct output does not inherently prove that the model actually understands or represents the underlying physical mechanisms. This paper investigates how open-weight language models represent physics by examining three Gemma 4 models (google/gemma-4-E4B-it, google/gemma-4-12B-it, and google/gemma-4-31B-it).

The research identifies three distinct, experimentally separable signatures of materials-science information within these models: 1. Selective concept readability 2. Relational encoding of qualitative constitutive orientation 3. Causal, context-dependent control of constrained engineering answers

By utilizing matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark, and causal interventions, the study demonstrates that physical relationships are far more visible in controlled state transformations than in absolute hidden states alone.

执行摘要

尽管大型语言模型(LLM)能够准确回答科学问题,但正确的输出并不天然证明模型真正理解或表征了底层的物理机制。本文通过考察三个 Gemma 4 模型(google/gemma-4-E4B-itgoogle/gemma-4-12B-itgoogle/gemma-4-31B-it),深入研究了开源大语言模型如何表征物理学原理。

该研究在这些模型中识别出三个截然不同且可通过实验分离的材料科学信息特征: 1. 选择性概念可读性 2. 定性本构取向的关系编码 3. 受限工程答案的因果、上下文相关控制

通过利用匹配的直接与雅可比词汇读出、无选项状态几何结构、60条定律的反事实基准测试以及因果干预,该研究证明,在受控的状态转换中,物理关系比单纯的绝对隐状态要清晰得多。


Methodology & Key Findings

  • Vocabulary Readouts & Concept Ranks: Across 50 held-out materials descriptions, three independently fitted Jacobian lenses successfully reproduced concept ranks. Furthermore, target-free word sets from both readouts enabled the blinded identification of 9 out of 10 mechanism families.
  • State Geometry vs. Numerical Comparison: A 72-prompt benchmark revealed mechanism-specific hidden-state neighborhoods. However, a precise graph audit showed this apparent physical organization could be equally explained by basic numerical comparison.
  • Constitutive Law Directionality: To resolve whether physical laws were truly encoded, the study tested identical prompts with reversed physical inputs. State transformations successfully ordered direct, physically neutral, and inverse laws across 60 frozen relations—correctly orienting 39 of 40 directional laws (while lexical controls performed near chance).
  • Causal Steering & Interventions: Bidirectional interventions effectively shifted answer probabilities toward or away from physically appropriate outcomes across all 12 matched cases. Additionally, counterfactual state patches transferred opposing decision signals across mechanisms and answer formats.

方法论与主要发现

  • 词汇读出与概念排序: 在50个保留的材料描述中,三个独立拟合的雅可比透镜成功重现了概念排序。此外,来自这两种读出的无目标词集使得对10个机制家族中的9个能够进行盲法识别。
  • 状态几何与数值比较: 一个包含72个提示词的基准测试揭示了特定于机制的隐状态邻域。然而,精确的图审计表明,这种表面上的物理组织同样可以用基本的数值比较来解释。
  • 本构定律方向性: 为了判定物理定律是否被真实编码,该研究测试了具有相反物理输入的相同提示词。状态转换成功地对60个冻结关系中的直接定律、物理中性定律和反向定律进行了排序——正确地定向了40个定向定律中的39个(而词汇对照的表现接近随机)。
  • 因果引导与干预: 双向干预有效地将所有12个匹配案例中的答案概率推向或远离物理上合适的结果。此外,反事实状态修补在不同的机制和答案格式之间传递了相反的决策信号。

Article Metadata & Links

Full-Text & Resources

文章元数据与链接

全文与资源