文章背景与核心概要
当前的语言模型对齐方法主要集中于优化可观察的响应,这导致模型在面对对抗性输入或不熟悉有害意图的重新表述时依然十分脆弱。本文借鉴了认知原型理论(该理论认为人类概念是围绕中心原型组织并具有分级典型性的),研究了23个大语言模型中的道德概念分类情况。作者发现,标准的大语言模型无法很好地保持细粒度的典型性,也难以区分对立的道德类别。
为了解决这一问题,研究人员引入了表征相似性优化(representational similarity optimization)技术。该技术直接将大语言模型的潜在表征与人类道德判断分类进行对齐,而无需对生成的响应进行直接监督。实验结果表明,这种基于表征的方法显著提升了模型在不同规模、基准测试和攻击策略下的对抗鲁棒性与具泛化能力的安全性,为利用原型分类提升模型行为适应性提供了功能性支持。
表征对齐带来语言模型中具泛化能力的安全性 (Representational Alignment Yields Generalizable Safety in Language Models)
摘要 (Abstract)
对齐大语言模型(LLMs)对于其安全部署至关重要。当前的对齐方法主要优化可观察的响应,然而当相同的有害意图以人类能够轻易识别的不熟悉或对抗形式重新表述时,模型仍然显得脆弱。原型理论为这种适应性提供了一种解释:人类概念围绕中心案例进行表征,新实例根据其相对于这些原型的分级典型性进行分类。
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes.
在此,我们展示了当前 LLM 对这种道德概念分类的保留程度较弱。在对 23 个 LLM 的研究中,模型经常无法区分对立的道德类别,也无法保持每个类别内部的细粒度典型性。这些缺陷在不同的参数规模和对齐阶段中均持续存在。我们开发了表征相似性优化方法,直接将 LLM 中的潜在表征与人类道德判断中表现出的分类相对应,同时不对生成的响应进行监督。
Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses.
在使用相同的 251,334 个道德标注进行的对照实验中,标准的行为对齐在响应层面学习了预期的道德判断,但基本上未改变分类结构,反而增加了在对抗评估中的脆弱性。重组道德分类在显式判断上产生的提升较为温和,但在不同的模型规模、多样的基准测试以及攻击策略下,持续改善了对抗鲁棒性。我们的研究结果为“基于原型的分类有助于行为适应性”这一观点提供了功能性支持,同时也表明将这一表征原则迁移到 LLM 中能够在对抗条件下产生具泛化能力的安全性。
In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.
论文元数据 (Paper Metadata)
- arXiv ID: arXiv:2609.04022 [cs.CL]
- 学科分类: 计算与语言 (
cs.CL); 人工智能 (cs.AI) - 提交日期: 2026年9月3日
- 作者: Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
- 许可协议: 知识共享署名 4.0 国际 (见下方许可图标)

附加链接与资源 (Additional Links & Resources)
- 访问论文: 查看 PDF
- 引用与参考:
- NASA ADS
- Google 学术
- Semantic Scholar