量子AI会梦见路径吗?基于Grover干涉的路径积分慢思考
文章背景与核心概要
本文探讨了将量子力学原理引入人工智能推理过程的前沿交叉领域。在大语言模型中,通过可验证奖励进行强化学习虽然能实现“慢思考”,但往往会导致“策略崩溃”(Policy Collapse)——即概率过度集中于一小部分成功的轨迹,从而侵蚀探索的多样性。
为了解决这一困境,研究人员提出了一种基于“路径积分慢思考”的新型量子AI框架。该方法将慢思考构建为推理轨迹上的相干量子动力学,使动作序列能够在叠加态中共存,并通过Grover振幅放大机制进行重新组合。通过解析训练目标和精确的状态向量模拟,该量子方法在保持测试集准确率和维持探索性路径多样性方面,显著优于传统的经典控制方法。
执行摘要 (Executive Summary)
While reinforcement learning with verifiable rewards allows large language models (LLMs) to engage in "slow thinking," it frequently leads to policy collapse—where probability concentrates exclusively on a small subset of successful trajectories, eroding exploratory diversity.
带可验证奖励的强化学习使大语言模型(LLM)能够进行“慢思考”,但这也会导致策略崩溃(policy collapse)——概率完全集中在一小部分成功的轨迹上,从而侵蚀了探索的多样性。
This paper investigates whether quantum AI can solve this dilemma through path-integral slow thinking. By formulating slow thinking as coherent quantum dynamics over reasoning trajectories, the authors demonstrate how action sequences can coexist in superposition and recombine via Grover amplitude amplification. Using an analytical training target and exact statevector simulations, the quantum approach significantly outperforms classical controls in preserving held-out accuracy and maintaining exploratory path diversity.
本文研究了量子AI是否能够通过路径积分慢思考来解决这一困境。通过将慢思考构建为推理轨迹上的相干量子动力学,作者展示了动作序列如何在叠加中共存,并通过Grover振幅放大进行重新组合。利用解析训练目标和精确的状态向量模拟,该量子方法在保持未见样本(held-out)准确率和维持探索性路径多样性方面,显著优于经典的控制方法。
摘要 (Abstract)
Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: probability concentrates onto a few successful trajectories and exploratory diversity erodes. We ask whether quantum AI can realize slow thinking differently.
带可验证奖励的强化学习使大语言模型能够进行慢思考,但同样的训练会引发策略崩溃:概率集中在少数成功的轨迹上,探索多样性随之侵蚀。我们探讨量子AI是否能以不同的方式实现慢思考。
We formulate slow thinking as coherent dynamics over reasoning trajectories, a discrete path integral in which action sequences coexist in superposition and recombine before measurement. In our trainable realization, an exact verifier partitions the ensemble into collective accepted and rejected components that interfere under Grover amplitude amplification. A finite Grover evolution is maximized when the pre-amplification success probability lies at an analytically determined value below one, so inference itself defines an interior training target and removes the monotonic pressure toward unit success.
我们将慢思考构想为推理轨迹上的相干动力学(即离散路径积分),其中动作序列在测量前以叠加态共存并重新组合。在我们可训练的实现中,精确验证器将整体划分为集体接受和拒绝的组件,它们在Grover振幅放大下产生干涉。当放大前的成功概率处于低于1的解析确定值时,有限的Grover演化达到最大化,因此推理本身定义了一个内部训练目标,并消除了趋向单位成功的单调压力。
In exact statevector simulations of a \(2\times 3\) sliding puzzle, Grover training reaches accuracy \(0.95\) on a \(32\)-question training set at one round, against \(0.73\) for the strongest classical control. On held-out questions specialization has a cost: an untrained uniform policy read out through the same amplification remains the strongest reference on this solution-dense benchmark, and quantum training preserves far more held-out accuracy than classical training—at four rounds with matched circuit applications the two quantum models reach \(3.2\) and \(3.9\) times the strongest classical controls. The number of training questions supported by fixed-size policies trained at each amplification budget also grows faster with the budget than with matched classical repetition.
在对 \(2\times 3\) 滑块拼图的精确状态向量模拟中,Grover训练在1轮内对32个问题的训练集达到了 \(0.95\) 的准确率,而最强的经典控制为 \(0.73\)。在未见过的测试问题上,专业化是有代价的:通过相同放大读出的未训练均匀策略在这个解密集的基准测试中仍然是最强参考,而量子训练比经典训练保留了更多的测试准确率——在具有匹配电路应用的四轮训练中,两个量子模型分别达到了最强经典控制的 \(3.2\) 倍和 \(3.9\) 倍。在每个放大预算下训练的固定大小策略所支持的训练问题数量,其增长速度也比匹配的经典重复更快。
These results establish a Grover-based realization of path-integral slow thinking: the interior target preserves exploratory path diversity, and ensemble-level interference converts it into verified performance.
这些结果确立了基于Grover的路径积分慢思考实现:内部目标保留了探索性的路径多样性,而集合水平的干涉将其转化为经过验证的性能。
核心亮点与发现 (Key Highlights & Findings)
- Path-Integral Slow Thinking: Replaces classical sequential sampling collapse with coherent superposition over multiple reasoning pathways.
- Grover Amplitude Amplification: Leverages exact verifiers to partition and interfere accepted/rejected path ensembles.
- Interior Training Target: Avoids monotonic pressure toward unit success by setting optimal pre-amplification success probabilities below one, preserving exploratory diversity.
- Empirical Performance (\(2\times 3\) Sliding Puzzle):
- Training Set: Reached \(0.95\) accuracy (vs. \(0.73\) for classical controls) in 1 round.
- Held-out Generalization: Retained \(3.2\times\) to \(3.9\times\) the accuracy of classical counterparts after 4 rounds of matched circuit applications.
- 路径积分慢思考: 用多个推理路径上的相干叠加代替了经典的顺序采样崩溃。
- Grover振幅放大: 利用精确验证器来划分和干扰接受/拒绝的路径集合。
- 内部训练目标: 通过将放大前的最佳成功概率设定在小于1的值,避免了朝向单位成功的单调压力,从而保留了探索性多样性。
- 实证表现(\(2\times 3\) 滑块拼图):
- 训练集: 在1轮内达到 \(0.95\) 的准确率(经典控制为 \(0.73\))。
- 未见样本泛化能力: 经过4轮匹配电路应用后,保留了经典对照组 \(3.2\) 到 \(3.9\) 倍的准确率。