从截断到承诺:均匀离散扩散模型中的持久上下文
文章背景与核心概要
本文由 Satoshi Hayakawa 撰写,深入研究了均匀状态离散扩散模型(uniform-state discrete diffusion models)。该类模型能够在保持所有位置均可修订的前提下并行更新所有 Token。为了理解从“临时的单步选择”向“永久的持久上下文”的转变,作者提出了承诺显示采样(Committed Reveal Sampling, CRS)——这是一种无需训练的采样器,能够存储选定的 argmax Token 并将其整合到后续的模型输入中。
通过理论分析和实证评估,本研究表明: * 在精确的前向过程中,选择干净 Token 的贝叶斯误差不会随着噪声的减小而增加。 * 保持所选 Token 的可见性,有助于并行预测在序列级别的选择上达成一致。 * 相比于标准的 top-\(p\) 截断和标量温度缩放基线,CRS 在生成困惑度(GenPPL)与熵之间实现了更好的权衡,这证明了支持集限制(support restriction)和持久上下文(persistent context)在生成建模中是两种截然不同的控制机制。
摘要 (Abstract)
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top-\(p\) rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. We ask what changes when selected hypotheses instead become persistent context for later predictions. We therefore propose committed reveal sampling (CRS), a training-free sampler that stores selected argmax tokens and inserts them into subsequent model inputs. Our analysis gives a rationale for selecting later and for keeping selected tokens visible. Under the exact forward process, the Bayes error of selecting a clean token cannot increase as noise decreases, while in a simple latent-mode model, keeping the selected token visible helps later parallel predictions agree on the same sequence-level choice. Empirically, paired experiments on Duo-distilled then separate this persistent effect from single-step top-\(p\) restriction and scalar temperature scaling. Under the same finalization rule, CRS without top-\(p\) truncation reaches lower generative perplexity (GenPPL) than fixed \(p=0.95\) and \(p=0.9\) baselines across budgets of 8--64 function evaluations (NFE). At 64 NFE, the comparison at matched unigram entropy also gives lower GenPPL for CRS, yielding a more favorable GenPPL--entropy tradeoff. Base Duo shows the same direction in a descriptive comparison, while other diversity and continuation metrics can rank these operating points differently. These results identify support restriction and persistent context as distinct controls of that tradeoff.
均匀状态离散扩散模型可以在保持所有位置均可修订的同时并行更新所有 Token。即使常用的 top-\(p\) 规则在某个位置只留下一个候选者,该选择也仅影响当前的逆向步骤,并可以在下一个采样步骤中被修改。我们探讨了当被选中的假设转变为后续预测的持久上下文时会发生什么变化。因此,我们提出了承诺显示采样(CRS),这是一种无需训练的 sampler,它存储选定的 argmax Token 并将其插入后续的模型输入中。我们的分析为稍后进行选择以及保持所选 Token 可见提供了理论依据。在精确的前向过程中,选择干净 Token 的贝叶斯误差不会随噪声减小而增加;而在一个简单的潜在模式模型中,保持所选 Token 可见有助于后面的并行预测在相同的序列级选择上达成一致。在实证方面,Duo-distilled 上的配对实验随后将这种持久效应与单步 top-\(p\) 限制和标量温度缩放区分开来。在相同的定型规则下,在 8–64 的函数评估次数(NFE)预算范围内,没有使用 top-\(p\) 截断的 CRS 比固定的 \(p=0.95\) 和 \(p=0.9\) 基线达到了更低的生成困惑度(GenPPL)。在 64 NFE 下,在匹配的一元语法熵(unigram entropy)下的比较也为 CRS 带来了更低的 GenPPL,从而产生了更有利的 GenPPL 与熵的权衡。Base Duo 在描述性比较中显示出相同的方向,而其他多样性和延续性指标可能会对这些工作点进行不同的排名。这些结果表明,支持集限制和持久上下文是该权衡中两种截然不同的控制手段。
访问与资源 (Access & Resources)
- 全文选项 (Full-Text Options):
- 查看 PDF (View PDF)
- HTML 版本 - 实验性 (HTML Version (Experimental))
- TeX 源码 (TeX Source)
- 外部书目工具 (External Bibliographic Tools):
- 谷歌学术 (Google Scholar)
- 语义学者 (Semantic Scholar)
- NASA ADS