跳转至

标准可解释模型:利用拉格朗日力学演绎设计可解释机器学习方法的通用理论

文章背景与核心概要

随着人工智能模型日益复杂,可解释性已成为调试、理解和控制模型输出的关键。然而,可解释机器学习领域长期缺乏用于演绎设计方法的通用理论,导致文献碎片化且评估协议不一致。

本文提出了“标准可解释模型”(Standard Interpretable Model, SIM),这是一种基于拉格朗日力学的通用理论。SIM 通过一组前提定义了目标用户的可解释性需求,并据此系统地推导出对称性和约束条件,从而构建出一个拉格朗日函数,其极小值点即对应最优的可解释模型。这种方法既可以通过更新不透明模型的参数来实现,也可以将约束直接编译进可解释的架构中。


论文元数据 (Paper Metadata)

字段 详情
arXiv ID arXiv:2606.12289 [cs.LG]
DOI 10.48550/arXiv.2606.12289
学科分类 机器学习 (cs.LG);人工智能 (cs.AI);神经与进化计算 (cs.NE)
提交历史 [v1] 2026年6月10日,周三
[v2] 2026年8月18日,周二 (当前版本)
许可协议 知识共享署名 4.0 国际 license icon

作者 (Authors)

  • Pietro Barbiero
  • Giovanni De Felice
  • Mateo Espinosa Zarlenga
  • Francesco Giannini
  • Filippo Bonchi
  • Mateja Jamnik
  • Giuseppe Marra
  • Ruggero Noris

摘要 (Abstract)

随着人工智能模型复杂度的增加,可解释性已成为理解、调试和控制其计算过程不可或缺的工具。然而,可解释性领域目前缺乏用于演绎设计可解释方法的通用理论。这种理论与方法之间的鸿沟导致了文献的碎片化和评估协议的不一致。

为了填补这一空白,我们引入了标准可解释模型(SIM),这是一种基于拉格朗日力学的通用理论,能够实现可解释方法的演绎设计。具体而言,SIM 通过一组前提条件总结了目标用户对可解释性的定义。基于这些前提,SIM 系统地推导出可解释性对称性及其相应的约束条件,进而塑造出一个拉格朗日函数的景观,其极小值点对应于最优的可解释模型。为了达到这些极小值,研究者既可以更新不透明模型的参数值以使其更具可解释性,也可以将约束条件编译进可解释的架构中。

我们通过实证表明,SIM 能够识别并解决现有方法(包括传统方法、基于概念的方法和机械可解释性方法)的局限性,突出了尚未被充分探索的研究方向,并为核心编程接口的设计提供了指导。SIM 的演绎性质不仅是一种研究方法,还为可解释性课程提供了教学基础,并可能改变科学界对这一长期碎片化学科的视角。

As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, interpretability lacks general theories to deductively design interpretable methods. This gap between theories and methods results in a fragmented literature and inconsistent evaluation protocols.

To fill this gap, we introduce the Standard Interpretable Model (SIM), a general theory grounded in Lagrangian mechanics that enables the deductive design of interpretable methods. Specifically, the SIM summarises, in a set of premises, what interpretability is for a target user. From these premises, the SIM systematically derives interpretability symmetries and corresponding constraints, which shape the landscape of a Lagrangian whose minima correspond to optimal interpretable models. To reach the minima, one can either update the parameter values of an opaque model to make it more interpretable or compile constraints into an interpretable architecture.

We empirically show that the SIM identifies and solves limitations of existing methods (including traditional, concept-based, and mechanistic interpretability), highlights underexplored research directions, and informs the design of core programming interfaces. Beyond being a research method, the deductive nature of the SIM offers pedagogical grounding for interpretability curricula and may shift the scientific community's perspective of a discipline that has long been fragmented.


获取全文与资源 (Access Full-Text & Resources)