跳转至

文章背景与核心概要

近年来,关于大语言模型(LLM)是否具备元认知能力——特别是检测并报告自身内部状态的能力——引发了广泛讨论。本文借鉴人类元认知研究的洞见,指出当前认为LLM具备内省能力的结论为时尚早。作者提出了验证真正模型内省所需的两个严格条件:特权访问(Privileged Access)与二阶计算(Second-Order Computation)。

通过在这些标准下重新评估现有的两类主要测试范式(隐藏状态标签预测和内部状态篡改检测),研究发现:模型的表现往往可以通过仅依赖输入特征的通用异常检测来解释,并没有真正展现出对内部表征的特权访问或元认知监控能力。这项研究已被COLM 2026接收,对当前关于LLM自我意识与自我监控的研究热潮敲响了理性的警钟。


Can LLMs Introspect? A Reality Check

Authors: Shashwat Singh, Tal Linzen, Shauli Ravfogel
Identifiers: arXiv:2605.26242 [cs.AI] | Accepted at: COLM 2026
Dates: Submitted on May 25, 2026; Last revised August 21, 2026
Links: View PDF | HTML Version | DOI


Executive Summary

Recent studies have claimed that Large Language Models (LLMs) possess the capacity for metacognition—specifically, the ability to detect and report their own internal states. Drawing on insights from human metacognition research, this paper argues that such conclusions are premature.

近期研究声称,大语言模型(LLM)具备元认知能力——具体而言,即检测并报告自身内部状态的能力。本文借鉴人类元认知研究的洞见,认为此类结论为时尚早。

The authors establish two strict conditions necessary to validate genuine model introspection: 1. Privileged Access: The testing paradigm must rely on internal states and cannot be solved solely via input-derived cues. 2. Second-Order Computation: The process must involve meta-representations of first-order, task-related representations, which requires experimental designs where first-order and second-order accounts yield divergent predictions.

作者确立了验证真正的模型内省所必需的两个严格条件: 1. 特权访问(Privileged Access): 测试范式必须依赖内部状态,而不能仅通过从输入中提取的线索来解决。 2. 二阶计算(Second-Order Computation): 该过程必须涉及对一阶、任务相关表征的元表征(meta-representations),这要求实验设计能使一阶和二阶解释产生不同的预测结果。

Re-evaluating two prominent testing paradigms under these criteria, the study finds that: * Hidden-State Label Prediction: Classifiers restricted to input-only cues match the models' in-context predictions, disproving privileged internal access. * Internal State Tampering Detection: Models fail to reliably distinguish between internal state interventions and input manipulations, indicating that their success stems from generic anomaly detection rather than genuine metacognitive monitoring.

在这些标准下重新评估两个著名的测试范式后,本研究发现: * 隐藏状态标签预测(Hidden-State Label Prediction): 仅受限于输入线索的分类器能够与模型的上下文预测相匹配,这反驳了模型对内部表征拥有特权访问的假设。 * 内部状态篡改检测(Internal State Tampering Detection): 模型无法可靠地区分内部状态干预与输入操纵,这表明它们的成功源于通用的异常检测,而非真正的元认知监控。

Ultimately, the authors conclude that current empirical evidence falls short of demonstrating metacognitive monitoring in LLMs.

最终,作者得出结论:当前的实证证据尚不足以证明大语言模型中存在元认知监控。


Abstract

Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, we argue that this conclusion may be premature. We identify two conditions that a paradigm needs to meet in order to establish introspection.

大语言模型能否检测并报告其自身的内部状态?许多近期研究认为它们可以。通过借鉴人类元认知研究的教训,我们认为这一结论可能为时尚早。我们识别出了一个范式为了确立内省所需满足的两个条件。

First, the test needs to require privileged access: it should not be solvable using cues available in the input. Second, it needs to require second-order computation: second-order, meta-representations of first-order, task-related representations. This condition cannot be satisfied by task performance alone: it requires designs under which second-order and first-order accounts make divergent predictions.

首先,测试需要具备特权访问:它不应仅通过利用输入中可用的线索来解决。其次,它需要涉及二阶计算:即对一阶、任务相关表征的二阶元表征。仅凭任务表现无法满足这一条件:它需要精妙的实验设计,在设计中一阶和二阶的解释会产生分歧的预测。

We re-examine two paradigms that have been used to argue for model introspection in light of these conditions: 1. Hidden-State Classification: Models must predict labels derived from their own hidden states; we find that classifiers that can only access the input match the models' in-context predictions, indicating that the original results do not demonstrate privileged access to internal representations. 2. State-Tampering Detection: Models must detect whether their internal states have been tampered with; we find they cannot reliably distinguish such interventions from manipulations of the input, suggesting that their success reflects generic anomaly detection rather than sensitivity to internal interventions in particular.

鉴于上述条件,我们重新审视了此前被用于论证模型内省的两类范式: 1. 隐藏状态分类(Hidden-State Classification): 模型必须预测由其自身隐藏状态导出的标签;我们发现,仅能访问输入的分类器与模型的上下文预测结果一致,这表明最初的结果并未证明模型对内部表征具有特权访问权限。 2. 状态篡改检测(State-Tampering Detection): 模型必须检测其内部状态是否遭到篡改;我们发现它们无法可靠地区分此类干预与对输入的操纵,这表明它们的成功反映的是通用的异常检测,而非对内部干预的特殊敏感性。

We conclude that current evidence is insufficient to establish metacognitive monitoring in LLMs.

我们得出结论:现有的证据尚不足以在LLM中确立元认知监控。


References & Associated Resources

(Note: Original article graphic/license asset reference preserved below)
license icon

(注:下方保留了原文章的图形/许可资产引用)
license icon