跳转至

文章背景与核心概要

高通量筛选(HTS)是早期药物发现的基石,然而由于单重复实验设计的限制,它经常面临极端的数据稀疏性问题。这一局限性阻碍了传统机器学习指标(如AUROC、灵敏度和特异性)的应用,因为这些指标通常需要较大的标记数据集来进行经验估计。

本文引入了一种新颖的基于模型的框架,直接从高通量筛选中使用的稳健效应量参数——严格标准化平均差(SSMD)中推导这些分类指标。通过假设高斯等方差,作者建立了SSMD与关键性能指标之间的闭式关系。这种方法使得即使在单重复场景下也能计算出精确的置信区间和估计量。与随着样本量增加可能会产生误导的经典统计效力不同,这种基于SSMD的方法提供了一种稳定、可解释的群间分离度量,有效地弥合了经典HTS统计学与机器学习评估理论之间的鸿沟。


Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays

Author: Xiaohua Douglas Zhang
Date: August 6, 2026
arXiv ID: 2608.07609
Subjects: Applications (stat.AP); Artificial Intelligence (cs.AI); Quantitative Methods (q-bio.QM); Machine Learning (stat.ML)


摘要

高通量筛选(HTS)是早期药物发现的基石,然而由于单重复实验设计的限制,它经常面临极端的数据稀疏性问题。这一局限性阻碍了传统机器学习指标(如AUROC、灵敏度和特异性)的应用,因为这些指标通常需要较大的标记数据集来进行经验估计。

High-throughput screening (HTS) is a cornerstone of early drug discovery, yet it frequently suffers from extreme data sparsity due to single-replicate experimental designs. This limitation hinders the use of traditional machine learning metrics—such as AUROC, sensitivity, and specificity—which typically require larger labeled datasets for empirical estimation.

本文引入了一种新颖的基于模型的框架,直接从高通量筛选中使用的稳健效应量参数——严格标准化平均差(SSMD)中推导这些分类指标。通过假设高斯等方差,作者建立了SSMD与关键性能指标之间的闭式关系。这种方法使得即使在单重复场景下也能计算出精确的置信区间和估计量。与随着样本量增加可能会产生误导的经典统计效力不同,这种基于SSMD的方法提供了一种稳定、可解释的群间分离度量,有效地弥合了经典HTS统计学与机器学习评估理论之间的鸿沟。

This paper introduces a novel, model-based framework that derives these classification metrics directly from the Strictly Standardized Mean Difference (SSMD), a robust effect-size parameter used in HTS. By assuming Gaussian equal-variance, the author establishes closed-form relationships between SSMD and key performance indicators. This approach allows for the calculation of exact confidence intervals and estimators even in single-replicate scenarios. Unlike classical statistical power, which can be misleading as sample sizes increase, this SSMD-derived method provides a stable, interpretable measure of group separation, effectively bridging the gap between classical HTS statistics and machine learning evaluation theory.


核心贡献

  • 指标推导:建立了SSMD与约登最优(Youden-optimal)灵敏度/特异性之间的数学桥梁。
  • 单重复效用:为超低重复筛选工作流中的分类性能评估提供了一种统计上严谨的方法。
  • 稳定性:提供了一种比经典统计效力更有意义的性能度量,因为它收敛于反映组间真实分离度的有限总体值。
  • 实证验证:使用丙型肝炎病毒初级siRNA筛选(约22,000个测量值)展示了该框架的有效性,证实了与传统的基于AUROC的方法相比,基于SSMD的阈值是等效的且高度可解释的。

Key Contributions

  • Metric Derivation: Establishes a mathematical bridge between SSMD and Youden-optimal sensitivity/specificity.
  • Single-Replicate Utility: Provides a statistically principled method for evaluating classification performance in ultra-low-replication screening workflows.
  • Stability: Offers a more meaningful performance measure than classical power, as it converges to a finite population value reflecting the true separation between groups.
  • Empirical Validation: Demonstrates the framework's effectiveness using a hepatitis C virus primary siRNA screen (approx. 22,000 measurements), confirming that SSMD-based thresholds are equivalent and highly interpretable compared to traditional AUROC-based methods.

元数据与访问

字段 详情
DOI https://doi.org/10.48550/arXiv.2608.07609
MSC 类别 92-08
ACM 类别 J.3
全文 查看 PDF

如需进一步探索,包括书目工具和相关研究,请参考原始的 arXiv 条目

Metadata & Access

Field Details
DOI https://doi.org/10.48550/arXiv.2608.07609
MSC Classes 92-08
ACM Classes J.3
Full Text View PDF

For further exploration, including bibliographic tools and related research, please refer to the original arXiv entry.