文章背景与核心概要
本文由 Piotr Jedryszek 和 Oliver M. Crook 撰写,探讨了如何利用稀疏自编码器(SAEs)等方法来解释生物学中的 AI 模型,并将其转化为科学发现的引擎。作者提出了一个集成流程,解决了潜在变量(latents)在可解释性上面临的核心挑战:一致性、可描述性以及预测能力。
在将该方法部署到 Boltz-1 Pairformer 主干网络时,他们的稳定性优先方法成功地识别出了具有可解释性的潜在变量,且评估次数显著减少,计算成本更低。这项研究为理解复杂生物学 AI 模型的内部机制提供了高效的工具,有助于推动AI驱动的科学突破。
Efficient Auto-Interpretability of AI Models in Biology
Summary
This paper, authored by Piotr Jedryszek and Oliver M. Crook, explores how methods like sparse autoencoders (SAEs) can be used to interpret AI models in biology and transform them into engines of scientific discovery. The authors propose an integrated pipeline that addresses the core challenges of latent interpretability: coherence, describability, and predictive power. When deployed on the Boltz-1 Pairformer trunk, their stability-prioritization method successfully identified interpretable latents with significantly fewer evaluations and lower computational costs.
本文由 Piotr Jedryszek 和 Oliver M. Crook 撰写,探讨了如何利用稀疏自编码器(SAEs)等方法来解释生物学中的 AI 模型,并将其转化为科学发现的引擎。作者提出了一个集成流程,解决了潜在变量可解释性的核心挑战:一致性、可描述性和预测能力。当部署在 Boltz-1 Pairformer 主干网络上时,他们的稳定性优先方法成功地识别出了可解释的潜在变量,并且评估次数明显减少,计算成本更低。
Metadata
- arXiv ID: arXiv:2608.27754 [q-bio.QM]
- Subjects: Quantitative Methods (
q-bio.QM); Artificial Intelligence (cs.AI) - Authors: Piotr Jedryszek, Oliver M. Crook
- Submitted: August 27, 2026
- License: Creative Commons Attribution 4.0 International

- arXiv ID: arXiv:2608.27754 [q-bio.QM]
- Subjects: Quantitative Methods (
q-bio.QM); Artificial Intelligence (cs.AI)- Authors: Piotr Jedryszek, Oliver M. Crook
- Submitted: August 27, 2026
- License: Creative Commons Attribution 4.0 International
Abstract
Sparse autoencoders (SAEs) and other interpretability methods could turn AI models in biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: 1. Whether it is coherent, 2. Whether it can be described, and 3. Whether that description has predictive power.
These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required:
- Cross-seed dictionary stability: Prioritises which latents are worth spending resources to investigate.
- Intruder-detection task: Asks whether a latent's activating examples share a recognizable pattern.
- Candidate biological description: Proposes a biological description in a separate pass, which we convert into falsifiable predictions that can be tested in silico.
When deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them. External checks show that the surfaced motifs are significantly enriched for their claimed annotations.
However, the results also suggest a possible tension: cross-seed stability might be selecting for some types of features (such as structure-related ones) much more than others (such as function-related features).
稀疏自编码器(SAEs)和其他可解释性方法通过解释生物学及其他领域中AI模型的超人能力,有望将其转化为科学发现的引擎。然而,只有当我们明确以下三点时,潜在变量才具有实际价值: 1. 它是否具有一致性; 2. 它是否可以被描述; 3. 该描述是否具有预测能力。
这些问题在常规研究中经常被混为一谈。我们将它们整合到一个单一的流程中,并报告了每个阶段所需的实践创新:
- 跨随机种子字典稳定性(Cross-seed dictionary stability): 优先确定哪些潜在变量值得投入资源进行研究。
- 入侵者检测任务(Intruder-detection task): 探究某个潜在变量的激活示例是否共享可识别的模式。
- 候选生物学描述(Candidate biological description): 在单独的遍历中提出生物学描述,并将其转化为可在计算机模拟(in silico)中测试的可证伪预测。
当部署在 Boltz-1 Pairformer 主干网络上时,稳定性优先方法能够找到可解释的潜在变量,每个变量所需的潜在评估次数减少了约 4.4 倍,测得的计算成本降低了 5.2 倍,同时恢复了一半以上的变量。外部检验表明,浮现出的基序在其声称的注释方面显著富集。
然而,结果也暗示了一种可能的张力:跨种子稳定性可能在很大程度上偏向于选择某些类型的特征(例如与结构相关的特征),而对其他特征(例如与功能相关的特征)的选择较少。
Access & Resources
- Full-Text Options:
- View PDF
- HTML (Experimental)
- TeX Source
- External References:
- NASA ADS
- Google Scholar
- Semantic Scholar
- Full-Text Options:
- View PDF
- HTML (Experimental)
- TeX Source
- External References:
- NASA ADS
- Google Scholar
- Semantic Scholar