扰动响应中证据、矛盾与脆弱性的分解
文章背景与核心概要
扰动方法广泛应用于解释机器学习模型的决策,其核心是通过测量输入改变时预测结果的变化来推断模型行为。然而,标准的响应幅度通常只能反映模型“反应了多少”,却无法捕捉到这种反应背后的“具体含义”。同样的响应幅度可能对应着支持事实与反事实的差异、相反的观点,或者在扰动路径中途激增但在终点处彻底消失。
为了解决这一局限性,本文引入了 DECAF(Evidence, Contradiction, And Fragility 的分解,即证据、矛盾与脆弱性分解)。DECAF 通过追踪成对输入在被逐步揭示时对比关系的发展,将模型响应拆解为三个可解释的组成部分:证据(\(E\))、矛盾(\(C\))和脆弱性(\(F\))。这种分解严格保持了常规响应的幅度(\(\text{Abs} = E + C + F\)),并且在终点相对公理(endpoint-relative axioms)下具有唯一性。
摘要
Perturbation methods are widely used to explain machine learning model decisions by measuring how predictions change when inputs are altered. However, standard response magnitude only reveals how much a model reacts, failing to capture the meaning behind that reaction. The same magnitude can support a factual-counterfactual difference, oppose it, or spike along the perturbation path only to vanish at the endpoint.
扰动方法广泛应用于解释机器学习模型的决策,其核心是通过测量输入改变时预测结果的变化来推断模型行为。然而,标准的响应幅度通常只能反映模型“反应了多少”,却无法捕捉到这种反应背后的“具体含义”。同样的响应幅度可能对应着支持事实与反事实的差异、相反的观点,或者在扰动路径中途激增但在终点处彻底消失。
To address this limitation, this paper introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility). DECAF tracks how contrasts develop as paired inputs are progressively revealed, splitting responses into three interpretable components: * Evidence (\(E\)) * Contradiction (\(C\)) * Fragility (\(F\))
为了解决这一局限性,本文引入了 DECAF(证据、矛盾与脆弱性分解)。DECAF 通过追踪成对输入在被逐步揭示时对比关系的发展,将响应拆解为三个可解释的组成部分: * 证据 (\(E\)) * 矛盾 (\(C\)) * 脆弱性 (\(F\))
This decomposition strictly preserves the ordinary response magnitude (\(\text{Abs} = E + C + F\)) and is unique under endpoint-relative axioms.
这种分解严格保持了常规响应的幅度(\(\text{Abs} = E + C + F\)),并且在终点相对公理下具有唯一性。
核心亮点与性能
Key Highlights & Performance
- Reliable Alignment: In a 72-model audit on ImageNet-9 featuring models with identical response magnitudes but divergent behaviors, the largest DECAF component correctly matched the observed behavior in 96.4% of cases (compared to just 35.0% for magnitude alone).
- 可靠的一致性: 在针对 ImageNet-9 的 72 个模型审计中(这些模型具有相同的响应幅度但行为各异),DECAF 的最大分量在 96.4% 的情况下能够正确匹配观察到的行为(相比之下,仅使用幅度的方法准确率仅为 35.0%)。
- Path Sensitivity: Changing the reveal path increases the total response by nearly 80%; while evidence remains virtually unchanged, fragility surges by more than \(4\times\).
- 路径敏感性: 改变输入揭示路径会使总响应增加近 80%;虽然证据几乎保持不变,但脆弱性却激增了 \(4\) 倍以上。
- Superior Benchmarking: On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform standard general-purpose attribution baselines.
- 卓越的基准测试表现: 在 FunnyBirds 和 ImageNet-1k 数据集上,简短的单向(前向)DECAF 轨迹表现优于标准的通用归因基线。
- Efficiency: On a 1B-scale DINOv2 model, a short DECAF trajectory matches a strong gradient-based baseline while achieving 4.75× lower wall-clock time and 2.36× lower peak memory.
- 高效性: 在 10 亿参数规模的 DINOv2 模型上,短轨迹的 DECAF 能够匹配强有力的基于梯度的基线,同时实现 4.75倍更低的挂钟运行时间 和 2.36倍更低的峰值内存消耗。
文档元数据
Document Metadata
- arXiv Identifier: arXiv:2608.12935 [cs.AI]
- Author: Lei You
- Primary Subject: Artificial Intelligence (
cs.AI)- Secondary Subjects: Machine Learning (
cs.LG)- Submitted: August 13, 2026
- DOI: 10.48550/arXiv.2608.12935
- License: Creative Commons Attribution 4.0 International
- arXiv 标识符: arXiv:2608.12935 [cs.AI]
- 作者: Lei You
- 主要学科: 人工智能 (
cs.AI) - 次要学科: 机器学习 (
cs.LG) - 提交时间: 2026年8月13日
- DOI: 10.48550/arXiv.2608.12935
- 许可证: 知识共享署名 4.0 国际许可协议

访问与资源
Access & Resources
- Full-Text Options:
- View PDF
- HTML Version (Experimental)
- TeX Source
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
- 外部引用与工具:
- 谷歌学术
- Semantic Scholar
- NASA ADS