跳转至

用机械可解释性解释内在的道德自我修正

文章背景与核心概要

大语言模型的“内在道德自我修正”(Intrinsic moral self-correction)是指大模型仅通过提示词(Prompting)就能精炼其道德判断或对齐输出的现象。尽管这一现象在各种任务中被广泛观察到,但其背后的底层运作机制此前尚不明确。本文作者提出假设:内在道德自我修正的运作原理是通过沿可解释的潜在方向引导隐藏表征来实现的。

通过对四个道德相关任务中的六个大语言模型进行评估,研究人员证明了由自我修正提示词诱发的表征偏移与对比引导向量(Contrastive steering vectors)高度契合。此外,即使这些引导向量是由完全不相交的语料库构建的,这种契合度依然保持稳健。至关重要的是,通过激活相加(Activation addition)应用这些偏移,能够比自我修正提示词和引导向量本身更有效地改变模型行为,这表明表征引导正是内在道德自我修正的核心机械驱动力。


Summary

Intrinsic moral self-correction refers to the phenomenon where a Large Language Model (LLM) refines its ethical judgments or aligns its outputs purely through prompting. Although widely observed across diverse tasks, the underlying mechanics of how this happens have remained unclear.

In this paper, the authors hypothesize that intrinsic moral self-correction functions by steering hidden representations along interpretable latent directions. Evaluating six LLMs across four morality-related tasks, the researchers demonstrate that representation shifts induced by self-correction prompts strongly align with contrastive steering vectors. Furthermore, this alignment remains robust even when the steering vectors are constructed from a completely disjoint corpus. Crucially, applying these shifts via activation addition can alter model behavior even more effectively than the self-correction prompts and steering vectors themselves, pointing to representation steering as the primary mechanistic driver of moral self-correction.


Paper Metadata

  • arXiv ID: arXiv:2505.11924 [cs.CL]
  • Authors: Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu
  • Primary Subject: Computation and Language (cs.CL)
  • Secondary Subjects: Artificial Intelligence (cs.AI), Machine Learning (cs.LG)
  • Journal Reference: 4th Deployable AI Workshop at AAAI 2026
  • Submission History:
  • [v1] Sat, 17 May 2025
  • [v2] Sun, 19 Oct 2025
  • [v3] Wed, 11 Feb 2026
  • [v4] Fri, 21 Aug 2026 (latest revision)

Abstract

Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across diverse tasks, its mechanism remains unclear. We hypothesize intrinsic moral self-correction functions by steering hidden representations along interpretable latent directions. Evaluating six LLMs across four morality-related tasks, we demonstrate that the representation shifts induced by self-correction prompts align with contrastive steering vectors. This alignment transfers even when the steering vectors are constructed from a disjoint corpus. Notably, when applied via activation addition, these prompt-induced shifts can alter model behavior more effectively than the self-correction prompts and the steering vectors. Our findings suggest representation steering is the mechanistic driver of intrinsic moral self-correction.