动态上下文调度:超越静态宇宙的学习
文章背景与核心概要
在上下文强化学习(Contextual RL)中,传统的训练方法通常假设环境参数在单个回合(episode)内部保持静态,而回合之间发生变化。然而,这与许多真实世界中参数随时间动态演变的部署场景相脱节。本文引入了一种名为“动态上下文调度”(Dynamic Context Scheduling)的全新训练方法,将回合内的上下文变化视为一种可控的塑形机制,而非单纯的部署现实。
通过引入预定的调度方案(如正弦偏移或余弦退火),该方法让策略在训练过程中接触到更丰富、更具时间结构的环境参数空间。作者开发了 DYNAMICCARLENV 框架,并在 CartPole、BipedalWalker 以及 VehicleRacing 等环境中进行了广泛实验。结果表明,动态调度在分布外(OOD)评估中持续匹配或超越了静态上下文基线,并且在复杂的 BipedalWalker 和 VehicleRacing 环境中实现了更高的分布内(ID)评估性能。此外,初步研究还表明,多阶段课程的自动搜索能够成功发现稳健的泛化调度方案,其效果可媲美对单阶段调度器进行的大规模网格搜索。
摘要 (Summary)
本文介绍了“动态上下文调度”,这是一种用于上下文强化学习的新型训练工具。我们没有将回合内的上下文变化仅仅视为一种部署现实,而是将其视为一种可控的塑形机制。因此,上下文在每个训练回合中根据预定的调度方案演变,从而使策略暴露于环境参数空间中更丰富、更具时间结构的的区域。我们引入了
DYNAMICCARNENV框架,该框架使用可插拔的调度族(例如正弦偏移或余弦退火)来包装上下文环境。在结合 CARL 上下文化技术的CartPole、BipedalWalker和VehicleRacing环境中,我们证明了动态调度在分布外(OOD)机制中与静态上下文基线相匹配或表现更佳。有趣的是,对于更复杂的BipedalWalker和VehicleRacing环境,我们也取得了更高的分布内(ID)评估性能。初步研究结果表明,自动搜索多阶段课程可以成功发现能够改善泛化能力的调度方案,其性能与对单阶段调度器进行的大规模网格搜索相当。This paper introduces Dynamic Context Scheduling, a novel training instrument for contextual reinforcement learning that treats intra-episode context variation as a controlled shaping mechanism rather than merely a deployment reality. By allowing the context to evolve within each training episode according to predetermined schedules (such as sinusoidal offsets or cosine annealing), policies are exposed to richer and more temporally structured regions of the environment parameter space.
The authors propose
DYNAMICCARLENV, a framework that wraps contextual environments using pluggable schedule families. Through experiments acrossCartPole,BipedalWalker, andVehicleRacingenvironments utilizing CARL contextualization, the dynamic schedules consistently match or outperform static context baselines in out-of-distribution (OOD) regimes. Furthermore, for more complex environments likeBipedalWalkerandVehicleRacing, the method achieves higher in-distribution (ID) evaluation performance. Preliminary findings also indicate that automated multi-stage curricula search can successfully discover robust generalization schedules comparable to extensive grid searches over single-stage schedulers.
论文元数据 (Paper Metadata)
- arXiv ID: arXiv:2608.20799 [cs.AI]
- 学科分类: 人工智能 (
cs.AI) - 作者: Martin Mráz, André Biedenkapp
- 提交时间: 2026年8月21日
访问与资源 (Access & Resources)
- 全文格式: 查看 PDF | HTML (实验性) | TeX 源码
- DOI: 10.48550/arXiv.2608.20799
- 外部引用与工具:
- Google Scholar
- Semantic Scholar
- NASA ADS