预测性记忆定位:从内部信号预测选择性干预路径
文章背景与核心概要
在大语言模型的研究中,激活引导(Activation steering)被广泛用于将语言模型局部的内部表征转化为可操作的控制方向。然而,仅仅做到“定位”并不意味着所选的方向具备选择性操作区间——即无法保证在实现目标修改的同时不引发能力退化或语义损坏。
为了解决这一局限性,本文作者引入了预测性记忆定位(Predictive Memory Localization, PML)框架。该框架将测得的网格干预路径建模为记忆定位的主要预测对象。PML 将经校准的目标移动与语义邻近破坏、能力损害进行解耦,利用低剂量因果响应来预测边际级别的选择性结果,从而指导风险感知的干预操作。
arXiv: 2608.12892 [cs.AI]
Authors: Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao
Submitted: 13 Aug 2026 (Last revised 14 Aug 2026)
License: CC BY 4.0 
arXiv: 2608.12892 [cs.AI]
Authors: Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao
Submitted: 13 Aug 2026 (Last revised 14 Aug 2026)
License: CC BY 4.0
Summary
Activation steering is widely used to translate localized internal representations of language models into actionable control directions. However, mere localization does not indicate whether a chosen direction possesses a selective operating regime—meaning it can achieve target changes without causing capability or semantic damage.
To address this limitation, the authors introduce Predictive Memory Localization (PML), a framework that models the measured-grid intervention path as the primary predictive object of memory localization. PML separates calibrated target movement from semantic-neighbor and capability damage, utilizing low-dose causal responses to forecast margin-level selective outcomes and guide risk-aware interventions.
Summary
Activation steering is widely used to translate localized internal representations of language models into actionable control directions. However, mere localization does not indicate whether a chosen direction possesses a selective operating regime—meaning it can achieve target changes without causing capability or semantic damage.
To address this limitation, the authors introduce Predictive Memory Localization (PML), a framework that models the measured-grid intervention path as the primary predictive object of memory localization. PML separates calibrated target movement from semantic-neighbor and capability damage, utilizing low-dose causal responses to forecast margin-level selective outcomes and guide risk-aware interventions.
Key Highlights & Findings
- Comprehensive Scale: Evaluated across a frozen study covering 3,000 records from 9 datasets and 14 domains, generating 30,000 distinct record-direction-layer paths and 210,000 path-strength evaluations.
- Geometric Performance: At layer 7, the geometry-derived RFM/AGOP direction achieves 13.1% target-any and 12.3% clean-any success rates, outperforming random baselines by 3.6 and 3.4 percentage points under record-paired bootstrap validation.
- Low-Dose Predictive Power: Across record-, dataset-, and domain-grouped splits, responses at a low dose of \(|\alpha| = 0.1\) serve as the strongest predictive signal for outcomes at higher disjoint strengths (\(|\alpha| \in \{0.25, 0.5\}\)).
- Risk-Aware Selectors: On held-out records, a predictor-driven selector adaptively chooses coefficients or abstains, improving utility while minimizing semantic-neighbor damage compared to train-tuned fixed-strength policies.
- Robustness: Across three residual-norm-matched base models, learned directions preserve selective-path gains, and low-dose responses yield a strong macro AUROC of 0.801–0.828 on held-out records.
Key Highlights & Findings
- Comprehensive Scale: Evaluated across a frozen study covering 3,000 records from 9 datasets and 14 domains, generating 30,000 distinct record-direction-layer paths and 210,000 path-strength evaluations.
- Geometric Performance: At layer 7, the geometry-derived RFM/AGOP direction achieves 13.1% target-any and 12.3% clean-any success rates, outperforming random baselines by 3.6 and 3.4 percentage points under record-paired bootstrap validation.
- Low-Dose Predictive Power: Across record-, dataset-, and domain-grouped splits, responses at a low dose of \(|\alpha| = 0.1\) serve as the strongest predictive signal for outcomes at higher disjoint strengths (\(|\alpha| \in \{0.25, 0.5\}\)).
- Risk-Aware Selectors: On held-out records, a predictor-driven selector adaptively chooses coefficients or abstains, improving utility while minimizing semantic-neighbor damage compared to train-tuned fixed-strength policies.
- Robustness: Across three residual-norm-matched base models, learned directions preserve selective-path gains, and low-dose responses yield a strong macro AUROC of 0.801–0.828 on held-out records.
Full-Text & External Links
- PDF: View PDF
- HTML: arXiv HTML (Experimental)
- Source Code & Data: Available via associated platforms including Hugging Face, CatalyzeX, and DagsHub.
- Citations & References: Google Scholar | Semantic Scholar | NASA ADS
Full-Text & External Links
- PDF: View PDF
- HTML: arXiv HTML (Experimental)
- Source Code & Data: Available via associated platforms including Hugging Face, CatalyzeX, and DagsHub.
- Citations & References: Google Scholar | Semantic Scholar | NASA ADS