跳转至

文章背景与核心概要

当前的动作条件视频生成模型通常受限于单一的机器人具身形态,这阻碍了它们利用包含丰富可泛化物理学习信号的大规模异构互联网视频数据。为了弥合这一差距,本文推出了 CLAP 框架,这是一种能够跨人类和机器人代理的多样化互联网规模视频进行训练的跨具身动作条件视频生成方法。

CLAP 的核心见解在于,无论执行者是谁,通用的物理定律都支配着时空动态。然而,跨具身学习并非易事,因为不同机器人平台的动作表征差异巨大,且人类视频通常完全缺乏动作标注。为此,CLAP 通过对齐异构动作空间(末端执行器位姿、语言指令和潜在动作)并采用基于课程的学习方案(先通过无标签视频学习基础物理先验,再落地到特定动作空间以实现零样本部署),成功构建了迄今为止最全面的动作条件视频世界模型套件,在多种机器人形态上达到了甚至超越了单具身基准模型的性能。


CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Summary

CLAP is a novel framework for cross-embodiment, action-conditioned video generation designed to learn generalizable physical laws from diverse, internet-scale human and robotic videos. By overcoming the challenge of disparate action spaces—using a combination of end-effector poses, language instructions, and latent actions—CLAP introduces a curriculum-based learning recipe. It first captures foundational physical priors from unlabeled videos using latent actions, then grounds them in specific action spaces for zero-shot deployment. Ultimately, CLAP delivers robust video world models that match or exceed single-embodiment baselines across various robotic morphologies (including DROID, Bridge, bimanual YAM robots, and G1 humanoids).

CLAP 是一个用于跨具身、动作条件视频生成的新型框架,旨在从多样化的互联网规模人类和机器人视频中学习可泛化的物理规律。通过克服异构动作空间的挑战(结合使用末端执行器位姿、语言指令和潜在动作),CLAP 引入了一种基于课程的学习方案。它首先利用潜在动作从无标签视频中捕捉基础物理先验,然后将其落地到特定的动作空间中以进行零样本部署。最终,CLAP 提供了强大的视频世界模型,在各种机器人形态(包括 DROID、Bridge、双臂 YAM 机器人和 G1 人形机器人)上达到了或超过了单具身基准模型。


Paper Overview

  • Authors: Kechen Liu, Ola Shorinwa
  • Primary Subject: Robotics (cs.RO)
  • Secondary Subjects: Artificial Intelligence (cs.AI), Computer Vision and Pattern Recognition (cs.CV)
  • arXiv Identifier: arXiv:2608.27406
  • Submission Date: August 27, 2026
  • Project Website: omni-clap.github.io

论文概览

  • 作者: Kechen Liu, Ola Shorinwa
  • 主要学科: 机器人学 (cs.RO)
  • 次要学科: 人工智能 (cs.AI)、计算机视觉与模式识别 (cs.CV)
  • arXiv 标识符: arXiv:2608.27406
  • 提交日期: 2026年8月27日
  • 项目网站: omni-clap.github.io

Abstract

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents.

CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions:

  1. Reconciling Disparate Action Spaces: CLAP aligns different action modalities using end-effector poses, language instructions, and latent actions.
  2. Curriculum-Based Learning Recipe: To resolve individual limitations, the framework first learns foundational physical priors across unlabeled video data using latent actions, and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks.

Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date—spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids).

摘要

当前最先进的动作条件视频模型通常受限于单一的机器人具身形态,这使得它们无法利用包含丰富物理学习信号的海量异构视频数据。为了弥合这一差距,我们推出了 CLAP,这是一个用于跨具身动作条件视频生成的框架,能够在人类和机器人代理的多样化互联网规模视频上进行训练。

CLAP 的立足点在于:无论执行者是谁,通用的物理定律都支配着时空动态。然而,跨具身学习并非易事,因为不同机器人平台的动作表征差异巨大,且人类视频通常完全缺乏动作标注。CLAP 通过以下核心贡献解决了这一根本挑战:

  1. 调和异构动作空间: CLAP 结合使用末端执行器位姿、语言指令和潜在动作,对齐了不同的动作模态。
  2. 基于课程的学习方案: 为了克服单一方法的局限性,该框架首先利用潜在动作在无标签视频数据上学习基础的物理先验,随后将其落地到末端执行器动作空间中,以实现对真实世界任务的零样本部署。

至关重要的是,CLAP 在 DROID 等具有挑战性的环境中接近或超越了最先进的单具身视频模型。这些性能优势通过少样本自适应进一步叠加,确立了训练单具身视频世界模型的新范式。最终,CLAP 提供了迄今为止最全面的动作条件视频世界模型套件——涵盖了多样化的动作条件空间(末端执行器、语言和潜在动作)以及各种机器人形态(包括跨具身、DROID、Bridge、双臂 YAM 机器人和 G1 人形机器人)。


访问链接