跳转至

文章背景与核心概要

近年来,围绕语言模型欺骗行为的研究和媒体报道频频将人类般的心理状态和能动性赋予AI系统。这种倾向极易危险地模糊“表面上的欺骗行为”与“实际底层欺骗机制”之间的界限。为了解决这种模糊性,本文引入了一个全面的因果分类法,用以区分先验承诺与回顾性报告、模型偏好与实现的输出、虚假偏好与对误导接收者效用的敏感性,以及欺骗行为与产生该行为的目标或策略的来源。

通过对两个开源模型家族进行受控的猜数字游戏和股票交易实验,作者证明了在缺乏特定内部机制的情况下,也可能出现看似具有欺骗性的行为。相反,其他干预措施则提供了直接的实证证据,表明接收者的信息状态能够因果性地影响欺骗性偏好。归根结底,虽然欺骗行为可以指向欺骗机制,但确立这种机制仍然不能证明模型在欺骗中具有真正的能动性。


从Deceptive Outputs到Deceptive Mechanisms:语言模型欺骗研究的因果框架

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

arXiv: 2609.04166 [cs.AI]
Submitted on: 3 September 2026
Author: Yakov Pyotr Shkolnikov
License: Creative Commons Attribution 4.0 license icon

arXiv: 2609.04166 [cs.AI]
Submitted on: 3 September 2026
Author: Yakov Pyotr Shkolnikov
License: Creative Commons Attribution 4.0 license icon


摘要

近期围绕语言模型欺骗(deception)的研究和媒体报道,经常将人类般的心理状态和能动性归因于AI系统。这种倾向极易危险地模糊表面上的欺骗行为与实际底层欺骗机制之间的界限。

为了解决这一模糊性,本文引入了一个全面的因果分类法(causal taxonomy),用以区分: * 先验承诺与回顾性报告 * 模型偏好与实际输出 * 虚假偏好与对误导接收者效用的敏感性 * 欺骗行为与产生该行为的目标或策略的来源

通过对两个开源模型家族进行受控的猜数字和股票交易实验,作者证明了在没有所提出的内部机制的情况下,也可能发生看似欺骗的行为。相反,其他干预措施提供了直接的实证证据,证明接收者的信息状态可以因果性地影响欺骗性偏好。最终,虽然欺骗行为可以指向欺骗机制,但确立这样的机制仍然不能证明模型在欺骗中具有真正的能动性。

Summary

Recent research and media coverage surrounding language-model deception frequently attribute human-like mental states and agency to AI systems. This tendency can dangerously blur the lines between mere deceptive-looking behavior and an actual underlying deceptive mechanism.

To address this ambiguity, this paper introduces a comprehensive causal taxonomy that distinguishes between: * Prior commitment versus retrospective report * Model preference versus realized output * False preference versus sensitivity to the utility of misleading a recipient * Deceptive behavior versus the provenance of the objective or strategy producing it

Through controlled guessing-game and stock-trading experiments across two open-weight model families, the author demonstrates that deceptive-looking behavior can occur in the absence of the proposed internal mechanism. Conversely, other interventions provide direct empirical evidence that a recipient's information state can causally influence deceptive preferences. Ultimately, while deceptive behavior can point toward a deceptive mechanism, establishing such a mechanism still does not prove genuine model agency in the deception.


元数据与链接

访问论文

Access Paper