跳转至

病理基础模型的分布鲁棒性边界

文章背景与核心概要

病理学基础模型(Pathology Foundation Models)在跨机构泛化时往往表现不佳,这主要是由组织制备、染色和扫描伪影等非生物学变异引发的“捷径学习(shortcut learning)”所导致的。尽管此前曾提出鲁棒性指数(Robustness Index, RI)来衡量局部表征几何结构是由生物学因素还是非生物学因素主导,但其构建方式存在结构性局限,导致跨模型比较不可靠。

为了解决这一问题,本文作者引入了跨混杂因素鲁棒性边界(Cross-confounder Robustness Margin, CRoMa)。这是一个带有正负号的、基于单样本的边界指标,用于评估:具有相同生物学特征但不同混杂因素的样本,是否比具有相同混杂因素但不同生物学特征的样本靠得更近。CRoMa 允许将鲁棒性分析为完整的分布,而不仅仅是一个单一的合并得分。在多个切片级和全切片级编码器上对 CRoMa 的评估表明,其具有一致的基准排名,并且更高的中位 CRoMa 与下游线性探测中减少的捷径诱导性能下降密切相关。

Pathology foundation models often struggle with generalisation across different institutions due to shortcut learning driven by non-biological variations (such as tissue preparation, staining, and scanning artifacts). While the Robustness Index (RI) was previously introduced to measure whether local representation geometry is dominated by biological or non-biological factors, its construction suffers from structural limitations that make cross-model comparisons unreliable.

To address this, the authors introduce the Cross-confounder Robustness Margin (CRoMa)—a signed, per-sample margin evaluating whether samples sharing biological traits but different confounders lie closer together than those sharing confounders but having different biology. CRoMa allows robustness to be analyzed as a full distribution rather than a single pooled score. Evaluating CRoMa across multiple tile-level and slide-level encoders revealed consistent benchmark rankings and demonstrated that higher median CRoMa correlates with reduced shortcut-induced performance drops in downstream linear probes.


Metadata & Document Information

属性 详情
arXiv ID arXiv:2607.25497 [cs.CV]
研究主题 计算机视觉与模式识别 (cs.CV); 人工智能 (cs.AI)
作者 Clément Grisi, Jeroen van der Laak, Geert Litjens
提交时间 2026年7月28日 (v1); 最新修订于 2026年8月21日 (v4)
许可协议 知识共享署名-相同方式共享 4.0 国际版
DOI 10.48550/arXiv.2607.25497
Attribute Details
arXiv ID arXiv:2607.25497 [cs.CV]
Subjects Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Authors Clément Grisi, Jeroen van der Laak, Geert Litjens
Submitted 28 July 2026 (v1); last revised 21 August 2026 (v4)
License Creative Commons Attribution-ShareAlike 4.0 International
DOI 10.48550/arXiv.2607.25497

Abstract

病理学基础模型编码了由组织制备、染色和扫描引入的非生物学变异,从而促成了捷径学习,破坏了跨机构的泛化能力。鲁棒性指数(RI)被提出来评估局部表征几何结构是由生物学变异还是非生物学变异主导。然而,其构建存在结构性局限,导致跨模型比较不可靠,因此迫切需要更具原则性的度量标准。我们引入了跨混杂因素鲁棒性边界(CRoMa),这是一个带有正负号的、按样本计算的边界,用于衡量具有相同生物学特征但不同混杂因素的样本是否比具有相同混杂因素但不同生物学特征的样本更接近。该指标针对每个样本进行定义,允许在相同队列上比较模型,并将鲁棒性分析为一种分布,而不是简化为单一的合并得分。我们在三个基准上的 20 个切片级编码器中评估了 CRoMa。基于中位 CRoMa 的模型排名在各个基准中表现出高度一致性(斯皮尔曼相关系数 rho ~ 0.90),但每个编码器仍然保留了由混杂因素主导的样本,其普遍性和严重程度存在显着差异。在单独基准上评估的四个全切片级编码器也出现了类似的模式,将分析扩展到了切片级表征之外。较高的中位 CRoMa 与下游线性探测中较小的捷径诱导性能损失相关,支持其作为对捷径敏感性的表征级指标。

Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions. The Robustness Index (RI) was proposed to assess whether local representation geometry is dominated by biological or non-biological variation. However, its construction suffers from structural limitations that make cross-model comparison unreliable, calling for a more principled metric. We introduce the Cross-confounder Robustness Margin (CRoMa), a signed, per-sample margin that measures whether samples sharing the same biology but different confounder lie closer than samples sharing the same confounder but different biology. It is defined for every sample, allowing models to be compared on the same cohort and robustness to be analysed as a distribution rather than reduced to a single pooled score. We evaluated CRoMa across 20 tile-level encoders on three benchmarks. Rankings by median CRoMa were highly consistent across benchmarks (Spearman rho ~ 0.90), yet every encoder retained confounder-dominated samples, whose prevalence and severity varied markedly. Similar patterns emerged for four slide-level encoders evaluated on a separate benchmark, extending the analysis beyond tile-level representations. Higher median CRoMa was associated with smaller shortcut-induced performance losses in downstream linear probes, supporting its use as a representation-level indicator of shortcut susceptibility.



保留自源文件的许可证图标引用: license icon

License Icon reference preserved from source: license icon