跳转至

为什么鲁棒性会减少叠加?

文章背景与核心概要

对抗样本的起源及其背后的机制一直是机器学习领域中活跃的研究方向。机械可解释性,尤其是神经网络中的“叠加(Superposition)”现象,为理解这一问题开辟了新的途径。此前,Gorton & Lewis (2025) 证明了对抗样本源于特征叠加,并从经验上表明对抗训练能够减少这种叠加,但他们并未给出这一现象背后的机械解释。

本文作者 Adam Elimadi 受 Ilyas 等人 (2019) 特征分类学的启发,提出了一种经验性解释。文章梳理出了一条明确的因果链:对抗训练会丢弃非鲁棒特征,从而导致模型需要表示的总特征数量减少,进而自然地降低了特征叠加的程度。该研究已被 COLM 2026 AI 可解释性研讨会(AIW)接受。


论文元数据

字段 详情
arXiv 标识符 arXiv:2608.22155 [cs.LG]
标题 Why Does Robustness Reduce Superposition?
作者 Adam Elimadi
提交时间 2026年8月23日
研究领域 机器学习 (cs.LG), 人工智能 (cs.AI)
备注 9页,4张图。已被 COLM 2026 AI 可解释性研讨会(AIW)接受
DOI 10.48550/arXiv.2608.22155

Paper Metadata

Field Details
arXiv Identifier arXiv:2608.22155 [cs.LG]
Title Why Does Robustness Reduce Superposition?
Author Adam Elimadi
Submitted August 23, 2026
Subjects Machine Learning (cs.LG), Artificial Intelligence (cs.AI)
Comments 9 pages, 4 figures. Accepted at the COLM 2026 Workshop on AI Interpretability (AIW)
DOI 10.48550/arXiv.2608.22155

摘要

对对抗样本及其起源的研究仍然是一个开放的研究领域。机械可解释性,尤其是叠加现象,为解决这一问题提供了新的途径。Gorton & Lewis (2025) 证明了对抗样本源于叠加,并通过实验表明对抗训练减少了叠加,但他们并未提供关于该现象为何发生的机械性解释。受 Ilyas 等人特征分类学的启发,我们提出了一种经验性解释,并追踪了以下因果链:对抗训练放弃了非鲁棒特征,导致需要表示的总特征减少,从而导致更少的叠加。

Abstract

The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adversarial training abandons non-robust features, leading to fewer total features to represent, resulting in less superposition.


额外资源与链接

license icon

license icon