滞后耦合:内部表示在具备因果效力之前就已经可读
文章背景与核心概要
本文研究了语言模型中内部表征在整个 Pythia 模型套件(从 160M 到 12B 参数,跨越 8 个检查点和 4 个任务族)上的发展轨迹。作者发现了一种被称为“滞后耦合”(lagged coupling)的现象,即内部表征通过线性探针(linear probes)被读取的时间,远早于它们对模型行为产生因果效力(causal efficacy)的时间。
核心发现表明,内部表征在早期检查点往往已经完全可读(AUROC \(\ge 0.990\)),但沿这些方向进行干预(steering)在很大程度上等效于无效。研究进一步揭示了三个可分离的轨迹:内部可读性、行为可读性和因果效力。这项预注册的研究通过 OLMo-2 的复现验证了这一方向性趋势,确立了语言模型可解释性中更广泛的发展瓶颈。
Summary
The paper investigates the developmental trajectory of internal representations in language models across the full Pythia suite (160M to 12B parameters, across eight checkpoints and four task families). The author discovers a phenomenon termed lagged coupling, wherein internal representations become readable by linear probes long before they achieve causal efficacy over the model's behavior.
Key takeaways include: * Internal Readability vs. Causal Efficacy: Internal representations are often fully readable (AUROC \(\ge 0.990\)) from the earliest checkpoints, whereas steering along those same directions remains largely null-equivalent. * Three Dissociable Tracks: 1. Internal readability (saturates early everywhere). 2. Behavioral readability (develops gradually and progressively later at larger scales). 3. Causal efficacy (predominantly null-equivalent, occasionally counterproductive early). * Representation Headroom: Representation headroom along the probe direction grows up to 57x with training and scale, while causal write-in stays minimal—indicating that variables are increasingly written into the representation yet increasingly ignored by the readout. * Replication: A pre-registered OLMo-2 replication preserves this directional trend at an attenuated magnitude, establishing a broader developmental bottleneck in language model interpretability.
元数据
Metadata
-
学科分类: 计算与语言 (
cs.CL);人工智能 (cs.AI)- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI)
- Subjects: Computation and Language (
-
备注说明: 15 页,5 个图表,7 个表格。针对完整 Pythia 套件的预注册发展可解释性研究,包含 OLMo-2 复现。
- Comments: 15 pages, 5 figures, 7 tables. Pre-registered developmental interpretability study on the full Pythia suite with an OLMo-2 replication.
-
DOI: 10.48550/arXiv.2609.01048
全文与资源
Full-Text & Resources
-
-
相关代码与工具: Hugging Face | alphaXiv | CatalyzeX | DagsHub
- Associated Code & Tools: Hugging Face | alphaXiv | CatalyzeX | DagsHub
-
引用文献: Google Scholar | Semantic Scholar | NASA ADS
- Citations: Google Scholar | Semantic Scholar | NASA ADS