跳转至

衡量大语言模型中的激活控制能力

文章背景与核心概要

随着大语言模型(LLM)能力的不断提升,安全部署日益依赖于潜在空间(latent-space)监测,以此来捕捉那些绕过标准行为评估的欺骗性或密谋性行为。然而,如果模型能够学会控制其自身的内部激活,就会产生一个关键的漏洞。

本文引入了激活可控性基准(Activation Controllability Benchmark),用于衡量LLM通过自然语言指令对其残差流激活进行调节的有效性。研究结果表明,在各种能力层级的大多数LLM都具备一定的时间维度控制能力,能够控制其残差流激活的方向和幅度。此外,这种内部控制能够(至少部分地)逃避标准的基于激活的监测工具(如线性探针、自然语言自编码器、激活神谕机和雅可比透镜)。作者建议,随着模型内省能力的提升,人工智能安全实验室和评估人员应在未来的前沿模型中积极追踪激活可控性。


摘要 (Abstract)

随着模型能力的不断增强,安全的部署很可能将依赖于潜在空间监测来作为行为评估的补充,尤其是当具备评估感知能力(evaluation-aware)的模型表现出阴谋诡计或欺骗行为时。然而,如果模型还能控制自身的激活,这种欺骗行为就可能延伸至潜在空间本身。基于这一考量,我们推出了激活可控性基准,以量化模型在多大程度上能够通过自然语言指令来调节其残差流。跨模型家族和能力水平的研究发现,大多数LLM能够以一定的时间分辨率控制其残差流激活的方向和幅度,尽管不同模型的表现存在显著差异。在简单任务中,这种程度的控制可以逃避基于激活的监测方法(包括线性探针、自然语言自编码器、激活神谕机和雅可比透镜),尽管这种逃避并非完美无缺。这些结果表明,随着内省能力的增强,对激活空间的控制可能会成为监测的一个混淆因素;因此,我们建议前沿实验室和评估人员在未来的模型中追踪激活可控性。

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.


论文元数据 (Paper Metadata)

  • arXiv ID: arXiv:2608.21664 [cs.AI]
  • 提交日期: 2026年8月21日
  • 学科分类: 人工智能 (cs.AI); 计算与语言 (cs.CL)
  • 作者:
  • Marek Mateusz Kowalski
  • Joshua Fonseca Rivera
  • Uzay Macar
  • David Demitri Africa
  • arXiv ID: arXiv:2608.21664 [cs.AI]
  • Submission Date: August 21, 2026
  • Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
  • Authors:
  • Marek Mateusz Kowalski
  • Joshua Fonseca Rivera
  • Uzay Macar
  • David Demitri Africa


文章许可证图标 (Article License Icon)

license icon

license icon