语言扩散模型是能够检索未见数据的联想记忆
文章背景与核心概要
本文研究了语言扩散模型(特别是基于均匀分布的离散扩散模型 UDDM)如何记忆训练数据并过渡到生成状态。通过将 UDDM 与联想记忆(AM)理论相联系,作者证明了可以通过条件似然最大化来形成吸引盆(basins of attraction),而无需显式的能量函数。研究揭示了一个由训练数据集大小调控的、从记忆到泛化的急剧转变过程,并且可以通过预测标记序列的条件熵来对其进行有效追踪。
语言扩散模型是能够检索未见数据的联想记忆
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
摘要
语言扩散模型在何时会死记硬背训练数据?我们又该如何定量评估其真实的生成状态?为了回答这些问题,本文指出基于均匀分布的离散扩散模型(UDDM)本质上表现为具有突现创造能力的联想记忆(AM)。联想记忆的核心思想是通过在存储的数据点周围建立不同的吸引盆,从而能够可靠地将这些数据点恢复为记忆。从历史上看,诸如霍普菲尔德网络(Hopfield networks)之类的模型使用显式的能量函数来保证这些稳定的吸引子。我们拓展了这一视角:利用能量并非绝对必要的观察,因为吸引盆同样可以通过条件似然最大化来形成。
通过评估训练集和测试集样本的标记恢复情况,我们发现 UDDM 存在一个由训练数据集大小所支配的、从记忆到泛化的急剧转变:随着数据集增大,训练样本周围的吸引盆收缩,而未见测试样本周围的吸引盆扩展,直到两者最终收敛到相同的水平。至关重要的是,我们仅通过预测标记序列的条件熵就能检测到这种转变:记忆阶段的特征是条件熵趋于消失,而在泛化状态下,大部分标记的条件熵仍然保持有限值。因此,条件熵为已部署模型的“记忆-泛化”转变提供了一种实用的探测手段。
Summary
This paper investigates how language diffusion models—specifically Uniform-based Discrete Diffusion Models (UDDMs)—memorize training data and transition into generative regimes. By connecting UDDMs to the theory of Associative Memories (AMs), the authors demonstrate that basins of attraction can be formed via conditional likelihood maximization without requiring an explicit energy function. The study reveals a sharp memorization-to-generalization transition modulated by training dataset size, which can be effectively tracked using the conditional entropy of predicted token sequences.
文档元数据
Document Metadata
| 字段 | 详情 |
|---|---|
| arXiv ID | arXiv:2604.26841 [cs.LG] |
| 作者 | Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri |
| 研究主题 | 机器学习 (cs.LG); 人工智能 (cs.AI); 计算与语言 (cs.CL) |
| 提交时间 | 2026年4月29日 (v1);最后修订:2026年9月2日 (v2) |
| 期刊引用 | 2026年自然语言处理实证方法会议论文集 (Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing) |
| DOI | 10.48550/arXiv.2604.26841 |
Field Details arXiv ID arXiv:2604.26841 [cs.LG] Authors Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri Subjects Machine Learning ( cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)Submitted 29 April 2026 (v1); Last revised: 2 September 2026 (v2) Journal Reference Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing DOI 10.48550/arXiv.2604.26841
摘要(原文)
Abstract
When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) with emergent creative capabilities. The core idea of an AM is to reliably recover stored data points as memories by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization.
By evaluating token recovery of training and test examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.
访问与资源
Access and Resources
- 全文选项: 查看 PDF | HTML (实验性) | TeX 源码
- 许可证: 知识共享署名 4.0
- 外部引用与工具:
- Google Scholar
- Semantic Scholar
- NASA ADS
- Hugging Face 模型与空间
- Full-Text Options: View PDF | HTML (Experimental) | TeX Source
- License: Creative Commons Attribution 4.0
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
- Hugging Face Models & Spaces