跳转至

离散扩散模型:从分词到生成的统一框架

文章背景与核心概要

离散去噪扩散模型(DDMs)作为处理离散数据的一种强大范式,正逐渐成为传统自回归(AR)模型的有力竞争者。与处理连续数据的扩散模型不同,离散扩散模型的核心在于如何构建离散状态空间,包括分词方案、词表拓扑结构以及特定领域的结构化字母表。

本文提出了一种统一的概念框架,旨在从底层离散状态空间设计的视角重新审视离散扩散模型。通过该框架,现有的多种方法(如转移矩阵法、掩码/吸收状态法以及基于分数/比率的方法)被整合为同一设计空间下的不同实例。此外,该研究还深入探讨了在训练目标、推理算法、扩展行为、系统优化及评估协议等方面的关键设计权衡,为未来的研究方向提供了重要指引。


摘要

离散去噪扩散模型(DDMs)最近已成为离散数据建模中自回归(AR)模型的一种极具吸引力的替代方案,它提供了并行生成和迭代全局优化的能力。与状态空间固定的连续扩散模型不同,DDMs 从根本上受到离散状态空间构建方式的影响:即分词方案、词表拓扑结构以及特定领域的结构化字母表。

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets.

本研究引入了一个统一的概念框架,通过底层离散状态空间的构建来审视离散扩散模型。在该框架内,现有的各种表述方式,包括转移矩阵法、掩码/吸收状态法以及基于分数/比率的方法,都表现为同一设计空间下的不同实例。该框架进一步揭示了在训练目标、推理算法、扩展行为、系统优化和评估协议等方面的共同设计权衡,并提出了几个有前景的未来研究方向。

This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.


文章元数据

字段 详情
arXiv ID arXiv:2607.13431 [cs.LG]
学科分类 机器学习 (cs.LG); 人工智能 (cs.AI); 计算与语言 (cs.CL)
提交历史 v1: 2026年7月15日
v2: 2026年8月25日 (当前版本)
DOI 10.48550/arXiv.2607.13431
许可协议 知识共享署名 4.0 国际许可协议

作者

Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Dawn Song, Philip S. Yu, Xue Liu


全文及资源链接

外部引用与工具


(许可图标参考占位符已从原文保留: license icon)