超越迁移准确率:面向低资源语言的机制引导受控自适应
文章背景与核心概要
在大型语言模型的跨语言迁移和低资源语言微调中,如何既提升目标语言的表现,又避免源语言及相关任务性能的灾难性遗忘,一直是AI研究的核心挑战之一。传统的模型机制发现(Circuit Discovery)方法通常依赖于具备清晰反事实(Counterfactuals)的模板化任务,这严重限制了其在多样化、非结构化自然文本上的应用。
为了突破这一瓶颈,本文作者针对非结构化文本场景,对 Transformer 上下文分解方法(CD-T)进行了创新性适配。通过引入标签平衡激活均值和任务方向相关性评分,该研究实现了无反事实(counterfactual-free)的电路发现。在此基础上,作者提出了“电路定向监督微调”(CT-SFT)策略,将参数更新严格限制在与任务相关的注意力头和 LayerNorm 层上。
实验证明,在 NusaX 跨语言情感迁移基准和 XNLI 任务上,CT-SFT 不仅在低资源适配方面极具竞争力,还能最稳定地避免灾难性遗忘,完美保留源语言与相关任务的性能。这项工作证实了电路定向自适应是一种比全局微调更具可控性、且有机制支撑的优选替代方案。
摘要 (Abstract)
Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformers (CD-T) for unstructured settings via label-balanced activation means and task-directional relevance scoring, enabling counterfactual-free circuit discovery. We leverage the discovered circuits for Circuit-Targeted Supervised Fine-Tuning (CT-SFT), restricting parameter updates to task-relevant heads and LayerNorm. Experiments on NusaX cross-lingual sentiment transfer show that CT-SFT is highly competitive for low-resource adaptation. While non-circuit sparse updates and full fine-tuning sometimes match target accuracy through capacity recruitment, CT-SFT most consistently avoids catastrophic forgetting, preserving source-language and related-task performance. Extensions to XNLI support the source-retention and intervention findings on a harder task and two model families, showing that circuit-targeted adaptation provides a more controlled, intervention-supported alternative to global fine-tuning.
现有的电路发现方法依赖于具有清晰反事实的模板化任务,这限制了它们在多样化自然文本上的使用。我们通过标签平衡激活均值和任务方向相关性评分,将 Transformer 上下文分解(CD-T)适配到非结构化设置中,从而实现了无反事实的电路发现。我们利用发现的电路进行电路定向监督微调(CT-SFT),将参数更新限制在任务相关的头和 LayerNorm 上。在 NusaX 跨语言情感迁移上的实验表明,CT-SFT 在低资源适配方面极具竞争力。虽然非电路稀疏更新和全量微调有时可以通过容量招募来匹配目标准确率,但 CT-SFT 最能一致地避免灾难性遗忘,同时保留源语言和相关任务的性能。对 XNLI 的扩展支持了在更难的任务和两个模型系列上的源保留和干预发现,表明电路定向自适应为全局微调提供了一种更可控、受干预支持的替代方案。
论文元数据 (Paper Metadata)
- arXiv Identifier: arXiv:2601.08146 [cs.CL] (v4)
- Primary Subject: Computation and Language (
cs.CL)- Secondary Subjects: Artificial Intelligence (
cs.AI); Machine Learning (cs.LG)- Authors: Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya
- Conference Status: Accepted as a Findings paper at EMNLP 2026
- Submission Dates: Submitted Jan 13, 2026; Last revised Sep 2, 2026.
- License: Creative Commons Attribution-ShareAlike 4.0 International
- arXiv 标识符: arXiv:2601.08146 [cs.CL] (v4)
- 主要学科: 计算与语言 (
cs.CL) - 次要学科: 人工智能 (
cs.AI); 机器学习 (cs.LG) - 作者: Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya
- 会议状态: 已被 EMNLP 2026 接收为 Findings 论文
- 提交日期: 2026年1月13日提交;最后修订于 2026年9月2日。
- 许可协议: 知识共享署名-相同方式共享 4.0 国际版

链接与资源 (Links & Resources)
- Full Text Access: View PDF | HTML (Experimental) | TeX Source
- Citation Tools: NASA ADS | Google Scholar | Semantic Scholar