跳转至

从生成到模拟:世界模型距离真正的模拟器还有多远?

文章背景与核心概要

随着扩散模型和大模型视频生成技术的迅猛发展,生成式世界模型正日益被视为传统模拟器(如物理引擎、游戏引擎以及强化学习环境)的潜在替代品。然而,单纯的“生成”与真正的“模拟”之间究竟存在多大的差距,目前仍缺乏系统的评估。

本文开展了一项基于能力的综合研究,通过引入由传统模拟器八大核心能力组成的外部标尺,对世界模型进行了全面评估。这八大能力包括:资产构建、物理引擎、交互性、可控性、稳定性、状态反馈、多样性以及评估指标。通过对2018年至2026年6月期间发表的200篇代表性工作进行梳理(涵盖潜在动力学、视频 generation/生成、联合嵌入预测三条主要技术路线),作者揭示了当前生成式世界模型的现状与局限性。


📌 执行摘要 (Executive Summary)

With the rapid advancement of diffusion models and large-scale video generation, generative world models are increasingly viewed as potential replacements for traditional simulators (such as physics engines, game engines, and reinforcement-learning environments). However, the exact gap between mere generation and true simulation remains systematically unassessed.

随着扩散模型和大尺度视频生成的迅猛发展,生成式世界模型越来越被视为传统模拟器(如物理引擎、游戏引擎和强化学习环境)的潜在替代品。然而,单纯生成与真正模拟之间的确切差距仍未得到系统的评估。

With the rapid advancement of diffusion models and large-scale video generation, generative world models are increasingly viewed as potential replacements for traditional simulators (such as physics engines, game engines, and reinforcement-learning environments). However, the exact gap between mere generation and true simulation remains systematically unassessed.

This paper presents a comprehensive, capability-based study evaluating world models against an external yardstick consisting of eight core capabilities of traditional simulators: 1. Asset construction 2. Physics engine 3. Interaction 4. Controllability 5. Stability 6. State feedback 7. Diversity 8. Evaluation metrics

本文提出了一项全面的、基于能力的研究,通过将世界模型与包含传统模拟器八大核心能力的外部标准进行对比来评估它们: 1. 资产构建 2. 物理引擎 3. 交互性 4. 可控性 5. 稳定性 6. 状态反馈 7. 多样性 8. 评估指标

This paper presents a comprehensive, capability-based study evaluating world models against an external yardstick consisting of eight core capabilities of traditional simulators: 1. Asset construction 2. Physics engine 3. Interaction 4. Controllability 5. Stability 6. State feedback 7. Diversity 8. Evaluation metrics

By mapping 200 representative works published between 2018 and June 2026 across three primary technical routes—latent dynamics, video generation, and joint-embedding prediction—the authors uncover key insights regarding the current state and limitations of generative world models.

通过梳理 2018 年至 2026 年 6 月期间发表的 200 篇代表性工作(横跨潜在动力学视频生成联合嵌入预测这三条主要技术路线),作者揭示了生成式世界模型当前状态与局限性的关键见解。

By mapping 200 representative works published between 2018 and June 2026 across three primary technical routes—latent dynamics, video generation, and joint-embedding prediction—the authors uncover key insights regarding the current state and limitations of generative world models.


🔍 关键发现与分析 (Key Findings & Analysis)

  • Functional Successes: World models have successfully achieved functional substitution in specific scenarios regarding interaction and controllability.
  • Current Shortcomings: They still fall short of traditional simulators in maintaining formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution.
  • The Most Neglected Flaw: State feedback is highlighted as the most critical cross-route shortcoming. Out of 163 implementation papers analyzed, only 6 expose a runtime interface for querying entity states or physical parameters.

  • 功能性成功: 世界模型已在关于交互性可控性的特定场景中成功实现了功能替代。

  • 当前短板: 在维持物理定律的正式保证、结构化状态反馈以及可复现的长视野演化方面,它们仍落后于传统模拟器。
  • 最被忽视的缺陷: 状态反馈被强调为跨技术路线最关键的短板。在分析的 163 篇实现论文中,仅有 6 篇公开了用于查询实体状态或物理参数的运行时接口。
  • Functional Successes: World models have successfully achieved functional substitution in specific scenarios regarding interaction and controllability.
  • Current Shortcomings: They still fall short of traditional simulators in maintaining formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution.
  • The Most Neglected Flaw: State feedback is highlighted as the most critical cross-route shortcoming. Out of 163 implementation papers analyzed, only 6 expose a runtime interface for querying entity states or physical parameters.

🚀 未来研究方向 (Future Research Directions)

The study outlines six essential research directions to bridging the gap between generation and true simulation: 1. Formalized Physics: Embedding strict physical laws rather than purely statistical approximations. 2. Unified Action Interface: Establishing standardized control protocols across models. 3. First-Class State Feedback: Designing models with explicit queryable states and physical parameters. 4. Long-Horizon Stability: Improving error accumulation and drift over extended execution horizons. 5. Downstream-Utility Evaluation: Benchmarking models based on task performance rather than visual fidelity alone. 6. Cross-Route Hybridization: Combining latent dynamics, video generation, and joint-embedding predictions to leverage their respective strengths.

该研究勾勒出了六个弥合生成与真正模拟之间差距的必经研究方向: 1. 形式化物理(Formalized Physics): 嵌入严格的物理定律,而非纯粹的统计近似。 2. 统一动作接口(Unified Action Interface): 在不同模型之间建立标准化的控制协议。 3. 一流的状态反馈(First-Class State Feedback): 设计具有显式可查询状态和物理参数的模型。 4. 长视野稳定性(Long-Horizon Stability): 改善在扩展执行视野下的误差积累与漂移。 5. 下游效用评估(Downstream-Utility Evaluation): 基于任务性能而非仅仅视觉保真度来对模型进行基准测试。 6. 跨路线混合(Cross-Route Hybridization): 结合潜在动力学、视频生成和联合嵌入预测,以发挥各自的优势。

The study outlines six essential research directions to bridge the gap between generation and true simulation: 1. Formalized Physics: Embedding strict physical laws rather than purely statistical approximations. 2. Unified Action Interface: Establishing standardized control protocols across models. 3. First-Class State Feedback: Designing models with explicit queryable states and physical parameters. 4. Long-Horizon Stability: Improving error accumulation and drift over extended execution horizons. 5. Downstream-Utility Evaluation: Benchmarking models based on task performance rather than visual fidelity alone. 6. Cross-Route Hybridization: Combining latent dynamics, video generation, and joint-embedding predictions to leverage their respective strengths.