跳转至

蛋白质结构预测:从进化约束到生成式建模

文章背景与核心概要

蛋白质结构预测是结构生物学的核心基础,决定了分子的功能并为机制解释提供了依据。近年来,深度学习的迅猛发展彻底改变了这一领域,使其从传统的、依赖多序列比对(MSA)的单体折叠,演变为能够对蛋白质复合物及日益复杂的异质分子系统进行建模的更广泛框架。

本文对蛋白质结构预测的方法论演进进行了系统性综述。作者从“表征与数据”、“架构与学习策略”、“置信度与评估”三个核心维度深入剖析了近期的技术突破,将该领域划分为四个不同的方法论阶段,并总结了三个跨维度的范式转变。这项工作为理解近期模型的演进路径、能力边界及实际应用提供了清晰的视角。


Summary

This review paper explores the methodological evolution of protein structure prediction, tracking its journey from traditional multiple sequence alignment (MSA)-driven monomer folding to advanced generative modeling frameworks. The authors examine recent breakthroughs through three core dimensions—representations and data, architectures and learning strategies, and confidence and evaluation—organizing the field into four distinct methodological phases and three cross-cutting paradigm shifts.

这篇综述论文探讨了蛋白质结构预测的方法论演进,追踪了它从传统的多序列比对(MSA)驱动的单体折叠,到先进的生成式建模框架的发展历程。作者通过三个核心维度(表征与数据、架构与学习策略、置信度与评估)检查了近期的突破,将该领域归纳为四个不同的方法论阶段和三个交叉的范式转变。


Metadata

  • arXiv ID: arXiv:2608.16094 [cs.AI]
  • Subject Areas: Artificial Intelligence (cs.AI), Machine Learning (cs.LG)
  • Submission Date: August 17, 2026
  • Authors: Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li
  • Comments: 15 pages, 4 figures, 4 tables. Preprint submitted to Elsevier
  • arXiv ID: arXiv:2608.16094 [cs.AI]
  • 学科领域: 人工智能 (cs.AI), 机器学习 (cs.LG)
  • 提交日期: 2026年8月17日
  • 作者: Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li
  • 备注: 15页,4张图表,4个表格。已向Elsevier提交预印本

Abstract

Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from multiple sequence alignment (MSA)-driven monomer folding into broader frameworks capable of modeling protein complexes and increasingly heterogeneous molecular systems.

准确的蛋白质结构预测是结构生物学的基石,因为蛋白质结构决定了其分子功能,并为机制解释提供了基础。深度学习的最新进展彻底改变了这一领域,使其从多序列比对(MSA)驱动的单体折叠,转变为能够对蛋白质复合物以及日益异质化的分子系统进行建模的更广泛框架。

Existing reviews have summarized this progress from the perspectives of representative models, application domains, and protein design. Building on these efforts, this review focuses on the methodological evolution of the field itself. It examines recent developments through three closely related dimensions: 1. Representations and data 2. Architectures and learning strategies 3. Confidence and evaluation

现有的综述已经从代表性模型、应用领域和蛋白质设计等角度总结了这一进展。在此基础上,本综述聚焦于该领域本身的方法论演进。它通过三个紧密相关的维度审视了近期的发展: 1. 表征与数据 2. 架构与学习策略 3. 置信度与评估

Within this perspective, the field is organized into four methodological phases and three cross-cutting transitions: * From explicit evolutionary coupling features and early contact prediction to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold. * From protein-only monomer folding to increasingly integrated modeling of heterogeneous molecular systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3. * From prediction-oriented structure inference to design-oriented generative modeling in RFdiffusion and related frameworks.

基于这一视角,该领域被组织为四个方法论阶段和三个交叉转变: * 从显式进化偶联特征和早期接触预测,到 AlphaFold2RoseTTAFoldESMFold 中学习到的序列表征。 * 从纯蛋白质单体折叠,到 AlphaFold-MultimerRoseTTAFoldNAAlphaFold3 中对异质分子系统日益一体化的建模。 * 从以预测为导向的结构推断,到 RFdiffusion 及相关框架中以设计为导向的生成式建模。

This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.

这一框架让我们能够更清晰地理解方法论的转变如何塑造了近期模型的能力、局限性以及实际作用。


Access & Resources