文章背景与核心概要
激活引导(Activation Steering)是一种在推理阶段通过向隐藏状态添加向量或特征来控制语言模型的技术。然而,这些引导信号的上游来源在过去常常被当作次要细节而忽视。
本文深入研究了“激活源选择”(Activation Source Selection)——即用于收集构建引导信号的隐藏状态的特定源上下文和激活读出策略的组合。通过对三个指令微调模型和四个引导任务族的评估,作者证明了仅改变源激活就会对引导成功率产生戏剧性的影响。研究表明,成功的激活引导从根本上取决于表示“模型即将做什么”,而不仅仅是已经出现了什么。
Where Steering Signals Come From: Activation Source Selection in Activation Steering
arXiv ID: arXiv:2607.25270
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Conference: Accepted to Findings of EMNLP 2026
Authors: Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
Submission Date: July 28, 2026 (Last revised August 28, 2026)
arXiv ID: arXiv:2607.25270
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Conference: Accepted to Findings of EMNLP 2026
Authors: Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
Submission Date: July 28, 2026 (Last revised August 28, 2026)
📌 Summary
Activation steering is a technique used to control language models at inference time by adding vectors or features to hidden states. However, the upstream source of these steering signals is often overlooked as a secondary detail.
This paper investigates activation source selection—the specific combination of source context and activation readout policies used to collect hidden states for building steering signals. Evaluating across three instruction-tuned models and four steering task families, the authors demonstrate that changing only the source activations drastically impacts steering success.
Key Findings & Contributions
- Execution-Boundary States Matter Most: Effective steering is not merely driven by whether a desired behavior appears in the source text. Instead, the strongest signals originate from execution-boundary states—the precise moments when a model is about to produce or continue the target behavior.
- The Pre-/Post-Realization Distinction: This insight explains why answer-based sources occasionally work: their useful components align with execution-boundary directions rather than simply relying on target text appearance.
- Tail Subtraction: Building on these findings, the authors introduce a new method called tail subtraction, which removes shared prompt and continuation semantics from boundary states, resulting in cleaner and more stable steering signals.
Overall, the research proves that successful activation steering depends fundamentally on representing what the model is about to do, rather than merely what has already appeared.
📌 摘要
激活引导是一种在推理阶段通过向隐藏状态添加向量或特征来控制语言模型的技术。然而,这些引导信号的上游来源往往被当作次要细节而忽视。
本文研究了激活源选择——即用于收集构建引导信号的隐藏状态的特定源上下文与激活读出策略的组合。通过对三个指令微调模型和四个引导任务族的评估,作者证明了仅仅改变源激活就会极大地影响引导的成功率。
核心发现与贡献
- 执行边界状态(Execution-Boundary States)最为关键: 有效的引导不仅仅由期望的行为是否出现在源文本中驱动。相反,最强的信号源自执行边界状态——即模型即将产生或延续目标行为的精确时刻。
- 实现前/实现后(Pre-/Post-Realization)的区别: 这一见解解释了为什么基于答案的源偶尔会起作用:它们的有用成分与执行边界方向保持一致,而不仅仅是依赖于目标文本的出现。
- 尾部相减法(Tail Subtraction): 基于这些发现,作者引入了一种名为尾部相减的新方法,该方法从边界状态中去除了共享的提示和延续语义,从而产生了更干净、更稳定的引导信号。
总而言之,该研究证明了成功的激活引导从根本上取决于表示模型即将做什么,而不仅仅是已经出现了什么。
📋 Metadata & Additional Links
- Full-Text Access: View PDF | HTML Version
- DOI: 10.48550/arXiv.2607.25270
- Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
📋 元数据与其他链接
- 全文访问: 查看 PDF | HTML 版本
- DOI: 10.48550/arXiv.2607.25270
- 引用与工具:
- Google Scholar
- Semantic Scholar
- NASA ADS