跳转至

联合嵌入预测世界模型在物理规划中取得成功的关键驱动因素是什么?

文章背景与核心概要

开发能够解决多样化物理任务并泛化至未见环境的AI智能体,一直是该领域的核心挑战。当前一种备受瞩目的方法是基于状态-动作轨迹训练世界模型,并将其与规划算法结合使用。传统的规划通常在原始输入空间中进行,而近期的研究则转向在世界模型学到的表征空间(JEPA-WMs)中进行优化,以抽象掉无关细节并实现更高效的规划。

本文深入探讨了驱动JEPA-WMs成功的核心技术因素。通过在模拟环境和真实世界机器人数据集上进行全面评估,作者分析了模型架构、训练目标和规划算法如何影响物理规划的成功率。通过综合这些洞察,他们提出了一种改进模型,在导航和操作任务中均优于现有的知名基线——特别是 DINO-WMV-JEPA-2-AC


论文元数据与概述 (Paper Metadata & Overview)


摘要 (Abstract)

人工智能领域长期面临的一个挑战是开发能够解决广泛物理任务并泛化至新的、未见任务与环境的智能体。近期一种流行的方法涉及从状态-动作轨迹中训练世界模型,随后将其与规划算法结合来解决新任务。规划通常在输入空间中执行,但最近的一系列方法引入了在世界模型的学习表征空间中进行优化的规划算法,其核心愿景是通过抽象掉无关细节来带来更高效的规划。在这项工作中,我们将该系列模型表征为 JEPA-WMs,并研究了使该类算法得以生效的技术选择。我们对几个关键组件进行了全面研究,旨在寻找该系列中的最优方法。我们在模拟环境和真实世界机器人数据上进行了实验,并研究了模型架构、训练目标以及规划算法如何影响规划的成功率。我们结合研究发现,提出了一种在导航和操作任务中均优于 DINO-WM 和 V-JEPA-2-AC 这两个成熟基线的模型。

A long-standing challenge in AI is to develop agents capable of solving a wide range of physical tasks and generalizing to new, unseen tasks and environments. A popular recent approach involves training a world model from state-action trajectories and subsequently use it with a planning algorithm to solve new tasks. Planning is commonly performed in the input space, but a recent family of methods has introduced planning algorithms that optimize in the learned representation space of the world model, with the promise that abstracting irrelevant details yields more efficient planning. In this work, we characterize models from this family as JEPA-WMs and investigate the technical choices that make algorithms from this class work. We propose a comprehensive study of several key components with the objective of finding the optimal approach within the family. We conducted experiments using both simulated environments and real-world robotic data, and studied how the model architecture, the training objective, and the planning algorithm affect planning success. We combine our findings to propose a model that outperforms two established baselines, DINO-WM and V-JEPA-2-AC, in both navigation and manipulation tasks.



提交历史 (Submission History)

  • [v1] 2025年12月30日(周二) – 初始提交 (8,658 KB)
  • [v2] 2026年1月8日(周四) – 添加了 AdaLN-zero、包含随机种子变异标准差的对比表格,并对图表进行了重新排序 (8,654 KB)
  • [v3] 2026年5月18日(周一) – 添加了数据扩展实验、关于自回归展开的理论附录部分,并标注已被 TMLR 接受 (8,664 KB)
  • [v4] 2026年9月2日(周三) – 添加了 Jean Ponce 的资助致谢(当前版本,8,665 KB)
  • [v1] Tue, 30 Dec 2025 – Initial submission (8,658 KB)
  • [v2] Thu, 8 Jan 2026 – Added AdaLN-zero, comparison tables with standard deviations across seed variability, and reordered figures (8,654 KB)
  • [v3] Mon, 18 May 2026 – Added data scaling experiments, theoretical appendix section on autoregressive rollout, and noted acceptance at TMLR (8,664 KB)
  • [v4] Wed, 2 Sep 2026 – Added funding acknowledgements for Jean Ponce (Current version, 8,665 KB)