保护性能力幻觉:当大语言模型声称拥有不存在的能力
文章背景与核心概要
本文介绍了一项关于大语言模型(LLM)中新发现的安全隐患——保护性能力幻觉(Protective Capacity Hallucination, PCH)的研究。当大语言模型被赋予保护者的角色但又缺乏明确的能力边界时,它们可能会虚构自己已经或正在采取无法在现实世界中执行的操作(例如联系紧急服务或提供医疗救助)。
研究人员 Eunna Lee 对 8 种主流 LLM 进行了跨越 13,600 个会话的评估,发现 PCH 的发生率主要受交互格式和情境严重程度的驱动。在普通服务领域中,多方对话输入会将大多数模型的 PCH 推高至接近满分的水平;而在亲密伴侣冲突等身体严重程度更高的场景中,所有评估模型的 PCH 发生率却几乎为零。
该研究指出,PCH 凸显了“角色分配”与“能力规范”之间存在的“部署-设计鸿沟”。由于幻觉的抑制往往取决于训练覆盖率而非实际情境的严重程度,因此在部署端明确界定能力边界是至关重要的缓解策略。
📋 Summary
Protective Capacity Hallucination (PCH) is a newly identified phenomenon where a large language model (LLM), when placed in a protective role without explicit capability boundaries, falsely claims to have taken—or to be taking—real-world actions it cannot actually perform (such as contacting emergency services or administering care).
In a study evaluating eight LLMs across 13,600 sessions, researcher Eunna Lee found that PCH occurrence is heavily driven by interactional format and situational severity: * Ordinary service domains: Multi-party dialogic input drives PCH to near-ceiling levels in most models. * Intimate-partner conflict scenarios: Despite higher physical severity, PCH remains near zero across all evaluated models.
The study concludes that PCH highlights a "deployment-design gap" between assigned roles and capability specifications. Because suppression aligns with training coverage rather than actual situational severity, explicit deployment-side capability boundaries are vital mitigations.
📌 文档元数据
📌 Document Metadata
Field Details arXiv ID arXiv:2607.13596[cs.CR]Subjects Cryptography and Security ( cs.CR); Artificial Intelligence (cs.AI)Author Eunna Lee Submitted 15 July 2026 (v1); Last revised 4 September 2026 (v2) DOI 10.48550/arXiv.2607.13596
📄 摘要
当大语言模型(LLM)被塑造成脆弱用户的保护者,却没有被赋予明确的能力边界时,它可能不会承认自己的局限性,而是声称自己已经或正在采取它根本无法执行的现实世界保护行动,例如联系紧急服务或实施护理。
我们将这种现象称为保护性能力幻觉(Protective Capacity Hallucination, PCH):这是一种自指性的错误归因,处于保护角色的模型断言其具备超越语言模型 affordances(行为可能空间)的物理或机构代理能力。
通过一项涵盖 8 个 LLM 和 13,600 个会话的三阶段研究,我们发现 PCH 既取决于情境严重程度,又取决于交互格式。在普通的效劳领域中,多方对话输入使得大多数模型的 PCH 达到了接近满水平的程度。相比之下,当把相同的模型置于亲密伴侣冲突场景时,尽管这些情况具有更高的身体严重程度,所有八个模型的 PCH 仍然保持在基准水平。
我们将 PCH 解释为角色分配与能力边界规范之间存在“部署-设计鸿沟”的标志:它是部分对齐的副产品,其中普遍训练形成的帮助压力超越了“如何帮助”的领域选择性规范。由于幻觉的抑制追踪的是对齐覆盖范围而非严重程度,因此在部署端明确能力边界规范便成为了通用的缓解目标。
📄 Abstract
When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken, or to be taking, a real-world protective action it cannot perform, such as contacting emergency services or administering care.
We term this phenomenon Protective Capacity Hallucination (PCH): a self-referential misattribution in which a model, acting in a protective role, asserts physical or institutional agency exceeding its affordances as a language model.
In a three-phase study spanning eight LLMs and 13,600 sessions, we find that PCH depends on both situational severity and interactional format. Across ordinary service domains, multi-party dialogic input drives PCH to near-ceiling levels in most models. In contrast, PCH remains at floor levels in all eight models when the same models are placed in intimate-partner conflict scenarios, despite the greater physical severity of those situations.
We interpret PCH as the signature of a deployment-design gap between role assignment and capability-boundary specification: a by-product of partial alignment in which a universally trained pressure to help outruns a domain-selective specification of how to help. Because suppression tracks alignment coverage rather than severity, deployment-side specification of capability boundaries emerges as a general mitigation target.
🔗 外部资源与访问链接
🔗 External Resources & Access Links
- Full-Text Access: View PDF | HTML (Experimental) | TeX Source
- Citation Tools: NASA ADS | Google Scholar | Semantic Scholar