言语过度自信的不同面相:一项可解释性研究
文章背景与核心概要
大语言模型(LLM)常常表现出过度自信的倾向,即在面对需要谨慎、保留意见或选择拒绝回答(abstention)的情形时,依然给出语气断然的回答。为了深入剖析这一现象,本文作者研究了 Qwen3-4B 模型在操纵逻辑必然性与可能性的受控推理场景下的表现,并重点考察了表达不确定性的三种模式:言语认知标记(verbal epistemic markers)、拒绝回答以及数字置信度分数。
研究结果揭示了驱动大模型过度自信的内部机制。通过模型可解释性技术,作者发现模型默认通过广泛的共享特征联盟来生成确定性,而不确定性则由一小部分专属特征介导的稀疏覆盖机制来实现。通过对这些不确定性特征进行干预,不仅证明了导致过度自信的潜在不对称性,还成功缓解了过度自信带来的错误。这组特征在不同的不确定性表达设置、多种语言以及分布外模态任务中均表现出极强的泛化能力,为改善模型校准与可靠性提供了重要的技术路径。
不同面相的言语过度自信:一项可解释性研究
Different Facets of Verbalised Overconfidence: an Interpretability Study
摘要
大型语言模型往往表现出过度自信,在证据表明应当采取保留态度或拒绝回答时仍给出确定的答案。通过利用操纵逻辑必然性和可能性的受控推理场景,我们在 Qwen3-4B 模型中研究了这一行为,考察了表达不确定性的三种方式:言语认知标记、拒绝回答以及数字置信度分数。
Abstract
Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores.
我们的结果证实了这种过度自信的倾向,特别是当模型被提示输出数字置信度分数时更为明显。在可解释性层面,我们提出了一种方法,能够差异化识别出负责不确定性和确定性的转码器(transcoder)特征。
Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty.
我们的分析表明,Qwen3-4B 的默认机制倾向于通过广泛的共享特征联盟来生成确定性,而不确定性则是通过由一小部分专用特征介导的稀疏覆盖(sparse override)来实现的。对这些不确定性特征进行干预,既从因果关系上证明了导致过度自信的这种失衡现象,也缓解了过度自信带来的错误。同一组特征在三种不确定性表达设置、多种语言以及分布外模态任务中都具有良好的泛化性。
Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigates overconfident errors. The same set of features generalizes across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.
核心发现与可解释性见解
Key Findings & Interpretability Insights
-
过度自信放大效应: 当模型被明确提示生成数字置信度分数时,其过度自信现象尤为显著。
- Overconfidence Amplification: Overconfidence is particularly pronounced when models are explicitly prompted to generate numeric confidence scores.
-
不对称机制:
- 确定性: 通过广泛的、默认的共享特征联盟来生成。
-
不确定性: 作为一种稀疏覆盖机制来实现,仅由一小部分专用的特征集合来介导。
- Asymmetric Mechanisms:
- Certainty: Generated via a broad, default coalition of shared features.
- Uncertainty: Implemented as a sparse override mechanism mediated by only a small, dedicated set of features.
-
因果干预: 直接干预这些稀疏的不确定性特征,成功证明了导致过度自信的潜在失衡,并减轻了错误的过度自信输出。
- Causal Intervention: Intervening directly on these sparse uncertainty features successfully proves the underlying imbalance causing overconfidence and mitigates erroneous overconfident outputs.
-
泛化能力: 所识别的不确定性特征具有鲁棒性,能够在不同的不确定性表达设置、多种语言以及分布外模态任务中无缝泛化。
- Generalizability: The identified uncertainty features are robust, generalizing seamlessly across different uncertainty-expression settings, multiple languages, and out-of-distribution modality tasks.
链接与全文获取
Links & Full-Text Access
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 外部学术工具: Google Scholar | Semantic Scholar | NASA ADS
- View PDF
- HTML Version (Experimental)
- TeX Source
- External Scholarly Tools: Google Scholar | Semantic Scholar | NASA ADS