文章背景与核心概要
本文研究了去中心化多智能体系统中,学习同伴(peers)如何为焦点智能体(focal agent)创造一个以智能体为中心的持续强化学习问题。当同伴更新其策略时,尽管全局博弈保持平稳,焦点智能体的奖励和环境动力学却会发生漂移。为了解决可复用行为结构在同伴更新下如何退化的核心问题,作者引入了“不变核心”(invariant core)的概念,即可在成功轨迹中频繁出现的最大抽象模式。
该研究的主要理论贡献是一个最坏情况下严格的条件定理(worst-case-tight conditioning theorem),证明了轨迹定律漂移(\(\varepsilon\))将覆盖率缩减限制在最多 \(\frac{\varepsilon}{p_0}\)。结合理论极限、走廊模型(\(R^2 > 0.9999\))以及在持续控制、cue-MNIST 和基于级别的觅食任务(Level-Based Foraging)中的实证研究,本文证明了核心侵蚀能够可靠地预测即将发生的失败,并支持接近先知(near-oracle)水平的干预措施。
巩固世界的边缘:多智能体世界边界中的持续学习问题 (Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary)
作者: Dane Malenfant
学科分类: 人工智能 (cs.AI)
arXiv: 2603.06813 [cs.AI]
DOI: 10.48550/arXiv.2603.06813
提交历史: 2026年3月6日提交;2026年8月24日最后修订。
摘要 (Summary)
本文研究了去中心化多智能体系统中的学习同伴如何为任何焦点智能体创造一个以智能体为中心的、由回合索引的诱导 MDP(Markov Decision Process)序列。联合博弈保持平稳,而焦点智能体的奖励和动力学发生漂移,从而形成了一个以智能体为中心的持续强化学习问题。将在回合内策略固定的同伴边缘化,可以保持每一个焦点轨迹定律和期望回报。因此,条件成功且可复用的结构在同伴更新下可能会发生退化。
This paper investigates how learning peers in a decentralized multi-agent system create an agent-centric continual reinforcement-learning problem. As peers update their policies, the focal agent's rewards and environment dynamics drift—even though the global game remains stationary.
不变核心(invariant core)通过高比例出现在成功焦点轨迹中的最大抽象模式来表示这种结构。主要结果是一个最坏情况下严格的条件定理:轨迹定律漂移 \(\varepsilon\) 最多会将候选对象条件成功的覆盖率降低 \(\frac{\varepsilon}{p_0}\),其中 \(p_0\) 是其参考成功质量(reference success mass),且该系数是精确的。同伴策略的运动提供了 \(\varepsilon\);正的覆盖率余量随后产生了一个经证实的 \(\Omega\left(\frac{1}{\eta}\right)\) 生存视界,并且在解析类中通过精确策略梯度实现的显式有效冲突条件下,产生了一个相匹配的 \(\Theta\left(\frac{1}{\eta}\right)\) 首次退出定律。通过校准成功质量和可执行性,相同的证书可以产生策略价值、库选择和传输遗憾保证。
To address how reusable behavioral structures degrade under peer updates, the author introduces the concept of an invariant core (maximal abstract patterns appearing frequently in successful trajectories). The primary theoretical contribution is a worst-case-tight conditioning theorem proving that trajectory-law drift (\(\varepsilon\)) bounds coverage reduction by at most \(\frac{\varepsilon}{p_0}\). Supported by theoretical limits, corridor models (\(R^2 > 0.9999\)), and empirical studies in continual control, cue-MNIST, and Level-Based Foraging, the work demonstrates that core erosion reliably predicts impending failure and enables near-oracle intervention.
一个可精确求解的走廊证实了结构预测,包括逆速寿命(\(R^2>0.9999\))。两项针对持续控制和 cue-MNIST 的注册 64 流研究表明,核心侵蚀可以预测即将到来的失败并实现近乎先知的干预;对八个已学习伙伴的 Level-Based Foraging 发展配对进行的探索性重新分析,暗示了在同伴学习下存在相同的“侵蚀-失败”联系。
In a stationary decentralized Markov game, learning peers generate an episode-indexed sequence of induced MDPs for any focal agent. The joint game remains stationary while the focal agent's rewards and dynamics drift, forming an agent-centric continual reinforcement-learning problem. Marginalizing peers whose policies are fixed within an episode preserves every focal trajectory law and expected return. Success-conditioned reusable structure may therefore degrade under peer updates.
An invariant core represents such structure through maximal abstract patterns appearing in a high fraction of successful focal trajectories. The main result is a worst-case-tight conditioning theorem: trajectory-law drift \(\varepsilon\) can reduce a candidate's success-conditioned coverage by at most \(\frac{\varepsilon}{p_0}\), where \(p_0\) is its reference success mass, and the coefficient is sharp. Peer-policy movement supplies \(\varepsilon\); positive coverage margin then yields a certified \(\Omega\left(\frac{1}{\eta}\right)\) survival horizon and, under an explicit effective-conflict condition realized by exact policy gradient in an analytic class, a matching \(\Theta\left(\frac{1}{\eta}\right)\) first-exit law. With calibrated success mass and executability, the same certificate yields policy-value, library-selection, and transfer-regret guarantees.
An exactly solvable corridor confirms the structural predictions, including the inverse-rate lifetime (\(R^2>0.9999\)). Two registered 64-stream studies in continual control and cue-MNIST show that core erosion predicts impending failure and enables near-oracle intervention; an exploratory reanalysis of eight learned-partner Level-Based Foraging development pairings suggests the same erosion–failure link under peer learning.
全文与资源 (Full-Text & Resources)
- PDF 文档: 查看 PDF
- HTML 版本: arXiv HTML (实验性)
- TeX 源码: arXiv e-Print 源码
- 许可协议: 知识共享署名 4.0 国际
- PDF: View PDF
- HTML Version: arXiv HTML (Experimental)
- TeX Source: arXiv e-Print Source
- License: Creative Commons Attribution 4.0 International
(注:根据明确的指令约束,在下方保留文章图片资源)
(Note: Preserved article image assets below per explicit instruction constraints)
