跳转至

文章背景与核心概要

有效的社会互动要求智能体能够同时将心理状态的推断转化为跨言语和非言语通道的协同行为信号。然而,现有的基准测试往往将心智理论(ToM)推理与具身行为孤立开来评估,从而在衡量社会推断如何转化为社会行动方面留下了关键的空白。

为了解决这一问题,作者推出了 MOSAIC(多模态社会行动、推断与通信编排,Multimodal Orchestration of Social Action, Inference, and Communication)。这是一个受控基准测试,用于评估在系统性变化的 ToM 约束下合作与竞争的场景。通过对 13 个模型(包括 11 个视觉语言模型,每个模型进行 200 次试验)的测试,该研究揭示了两个主要发现:1. 无法转化信念: VLM 未能在特定的 ToM 阶数约束下产生与预期结果一致的行为,并且施加显式约束也不会产生与指定推理水平相一致的可靠行为变化。2. 顺序瓶颈: 大多数模型难以产生方向一致的非言语信号。即使存在此类信号,VLM 智能体也无法解释并适当地对其他智能体的行为做出反应。

相反,作为带有显式 ToM 模块的结构化架构参考点的 PCM-LLM 在所有条件下均取得了成功,这表明显式信念-行动耦合足以克服这些局限性。


Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

arXiv: 2608.20975 [cs.AI]
Submitted: August 21, 2026
Authors: Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf


📌 Summary

Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Current benchmarks often evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving a critical gap in measuring how social inference translates into social action.

To address this, the authors introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark evaluating cooperative and competitive scenarios under systematically varied ToM constraints. Testing 13 models (including 11 Vision-Language Models) across 200 trials each, the study uncovers two primary findings: 1. Failure to Translate Beliefs: VLMs fail to produce behaviors consistent with expected outcomes under specific ToM-order constraints, and imposing explicit constraints yields no reliable behavioral change aligned with the specified reasoning level. 2. Sequential Bottlenecks: Most models struggle to produce directionally coherent nonverbal signals. Even when such signals are present, VLM agents fail to interpret and react appropriately to the behaviors of others.

Conversely, PCM-LLM—included as a structured architectural reference point featuring an explicit ToM module—succeeded across all conditions, indicating that explicit belief-action coupling is sufficient to overcome these limitations.


Metadata Attribute Details
Primary Subject Artificial Intelligence (cs.AI)
Cite As arXiv:2608.20975 [cs.AI]
DOI 10.48550/arXiv.2608.20975
License Creative Commons Attribution 4.0 International (license icon view license)

Access Options


🔍 Abstract

Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.


🔗 External Resources & Tools