跳转至

文章背景与核心概要

本文提出了一种强化学习状态占用度量(occupancy measure)的替代性表征方法,通过将规划准则嵌入到包含重置(resetting)机制的动力学过程中实现。研究所得的平稳测度(称为“访问测度”,visitation measure)为决策的信息几何提供了一个自然的数学基础。

作者证明了可实现的访问测度构成了一个双平坦(dually flat)的统计流形,其中访问概率和对数策略(log-policies)扮演了由条件熵连接的对偶仿射图表(dual affine charts)的角色。该几何框架将“规划即推理”(planning-as-inference)从线性奖励推广到处理访问测度的非线性泛函,每个迭代步均可通过一步自然梯度(natural-gradient)求解,同时为时序差分(temporal-difference)误差赋予了边际效用估计的新解释。这项研究深入探讨了其几何结构及其对强化学习和理论神经科学的重要意义。


规划即推理的双平坦几何

摘要 (Summary)

This paper introduces an alternative characterization of the reinforcement learning occupancy measure by embedding the planning criterion into the dynamics via a resetting planning process. The resulting stationary measure—the visitation measure—serves as a natural foundation for the information geometry of decision-making. The authors demonstrate that achievable visitation measures form a dually flat statistical manifold, where visitation probabilities and log-policies act as dual affine charts connected by conditional entropy. This geometric framework generalizes planning-as-inference to handle nonlinear visitation functionals solved via natural-gradient steps, while providing a novel interpretation of the temporal-difference error as a marginal-utility estimate.

本文引入了一种强化学习占用度量的替代表征,通过重置规划过程将规划准则嵌入到动力学中。由此产生的平稳测度——即访问测度(visitation measure)——构成了决策信息几何最自然的基石。作者证明,可实现的访问测度形成了一个双平坦统计流形,其中访问概率和对数策略充当由条件熵连接的双重仿射图表。这一几何框架将规划即推理从线性奖励推广到处理访问的非线性泛函,每个迭代步骤通过单步自然梯度求解,并赋予了时序差分误差作为边际效用估计的新解释。


论文元数据 (Paper Metadata)

作者 (Authors)

  • Nikola Milosevic
  • Asaki Kataoka
  • Nicolas Hinrichs
  • Kenji Doya
  • Nico Scherf

摘要原文 (Abstract)

We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.

我们提出了强化学习占用度量的一种替代表征,通过重置规划过程将规划准则嵌入到动力学中获得。其平稳测度被我们称为访问测度,是表达决策信息几何最自然的客体。可实现的访问测度形成了一个双平坦统计流形,其两个仿射图表分别是访问概率和对数策略,它们在条件熵下是对偶的。这种结构使得规划即推理能够从线性奖励推广到访问的非线性泛函,每个迭代通过一个自然梯度步骤求解,并赋予时序差分误差以边际效用估计的解释。我们发展了这一几何及其对强化学习和理论神经科学的影响。



外部参考与工具 (External References & Tools)