为什么鲁棒性会减少叠加?
文章背景与核心概要
对抗样本的起源及其背后的机制一直是机器学习领域中活跃的研究方向。机械可解释性,尤其是神经网络中的“叠加(Superposition)”现象,为理解这一问题开辟了新的途径。此前,Gorton & Lewis (2025) 证明了对抗样本源于特征叠加,并从经验上表明对抗训练能够减少这种叠加,但他们并未给出这一现象背后的机械解释。
本文作者 Adam Elimadi 受 Ilyas 等人 (2019) 特征分类学的启发,提出了一种经验性解释。文章梳理出了一条明确的因果链:对抗训练会丢弃非鲁棒特征,从而导致模型需要表示的总特征数量减少,进而自然地降低了特征叠加的程度。该研究已被 COLM 2026 AI 可解释性研讨会(AIW)接受。
论文元数据
| 字段 | 详情 |
|---|---|
| arXiv 标识符 | arXiv:2608.22155 [cs.LG] |
| 标题 | Why Does Robustness Reduce Superposition? |
| 作者 | Adam Elimadi |
| 提交时间 | 2026年8月23日 |
| 研究领域 | 机器学习 (cs.LG), 人工智能 (cs.AI) |
| 备注 | 9页,4张图。已被 COLM 2026 AI 可解释性研讨会(AIW)接受 |
| DOI | 10.48550/arXiv.2608.22155 |
Paper Metadata
Field Details arXiv Identifier arXiv:2608.22155 [cs.LG] Title Why Does Robustness Reduce Superposition? Author Adam Elimadi Submitted August 23, 2026 Subjects Machine Learning ( cs.LG), Artificial Intelligence (cs.AI)Comments 9 pages, 4 figures. Accepted at the COLM 2026 Workshop on AI Interpretability (AIW) DOI 10.48550/arXiv.2608.22155
摘要
对对抗样本及其起源的研究仍然是一个开放的研究领域。机械可解释性,尤其是叠加现象,为解决这一问题提供了新的途径。Gorton & Lewis (2025) 证明了对抗样本源于叠加,并通过实验表明对抗训练减少了叠加,但他们并未提供关于该现象为何发生的机械性解释。受 Ilyas 等人特征分类学的启发,我们提出了一种经验性解释,并追踪了以下因果链:对抗训练放弃了非鲁棒特征,导致需要表示的总特征减少,从而导致更少的叠加。
Abstract
The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adversarial training abandons non-robust features, leading to fewer total features to represent, resulting in less superposition.
额外资源与链接
- 全文访问:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 许可协议: 知识共享署名 4.0 国际许可协议
- 外部引用与工具:
- 谷歌学术
- Semantic Scholar
- NASA ADS
Additional Resources & Links
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution 4.0 International
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
