跳转至

激活引导能否捕捉多维度的作者风格?

文章背景与核心概要

本文探讨了激活引导(Activation Steering)技术是否能够有效捕捉大语言模型(LLM)中作者风格(authorship style)的多维且复杂的本质。作者们引入了面向方面的激活引导(Aspect-Aware Activation Steering, A3S)这一无需训练的框架,该框架通过结合干扰感知聚合(interference-aware aggregation)和实例特定的强度调优(instance-specific strength tuning)来合并各个方面的对比方向。

A3S 成功增强了多方面基准测试中的作者风格迁移,其性能超越了经过训练的基线模型,并保持了极小的目标与样本重叠度。这项研究为在无需自然语言风格描述符或专用训练的情况下,直接在激活空间中构建丰富的风格表示提供了新的途径。


摘要

Abstract

激活引导在控制大语言模型生成具有明确定义的属性方面展现出了良好的前景,但目前尚不清楚它是否能够处理作者风格这种多维度且难以定义的特性。我们探讨了是否可以通过沿修辞动机维度的结构化对比提示,直接在激活空间中构建丰富的风格表示,从而避开对自然语言风格描述符或专用训练的需求。

Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it remains unclear whether it can handle the multidimensional and hard-to-define nature of authorship style. We ask whether structured contrastive prompting along rhetorically-motivated dimensions can construct rich style representations directly in activation space, bypassing the need for natural language style descriptors or dedicated training.

我们发现,由此产生的方向共享一个共同的作者风格主干(backbone),但在承载真实风格信号的方面特定残差(aspect-specific residuals)上存在冲突,这解释了为什么简单的聚合方法会失败。我们将这一发现付诸实践,提出了面向方面的激活引导(Aspect-Aware Activation Steering, A3S),这是一个无需训练的框架,它通过干扰感知聚合来合并各个方面的对比方向,并针对每个实例调整引导强度。A3S 在真正具备多方面的作者风格迁移任务中表现更佳,在分布外基准测试的偏好评估中优异于经过训练的基线,并使目标-样本重叠度保持在较低水平。

We find that the resulting directions share a common authorship backbone while conflicting on aspect-specific residuals that carry genuine stylistic signal, explaining why naive aggregation fails. We operationalize this in Aspect-Aware Activation Steering (A3S), a training-free framework that merges per-aspect contrastive directions with interference-aware aggregation and tunes steering strength per instance. A3S improves authorship style transfer where it is genuinely multi-aspect, outperforms a trained baseline in preference evaluations on out-of-domain benchmarks, and keeps target-exemplar overlap consistently low.


访问与资源

Access & Resources