文章背景与核心概要
传统的虚拟代理非言语行为生成系统主要聚焦于“对齐”——将话语转化为单纯强调或图解所说内容的姿态和表情。然而,人类的沟通要复杂得多,非言语动作受到人际关系、说话者角色、社会背景以及内部情绪状态的深刻影响。因此,非言语线索可以强化、削弱、限定甚至完全矛盾于言语信息,同时也可能暴露出隐藏或偶然的内部状态(如情绪“泄漏”)。为了构建更逼真、更具人性化的虚拟代理(特别是针对咨询模拟等敏感训练场景),研究人员亟需能够捕捉这种微妙相互作用的模型。
本文借鉴埃克曼(Ekman)关于言语与非言语关系的理论框架,提出了一种行为不匹配的分类体系。研究探讨了大语言模型(LLM)是否能根据对话交互,有效选择出符合语境的、不一致的言语和非言语行为,并通过人类被试实验对这些具身行为进行了全面评估。
LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans
| Metadata | Details |
|---|---|
| arXiv ID | arXiv:2608.22731 |
| Primary Subject | Artificial Intelligence (cs.AI) |
| Secondary Subjects | Human-Computer Interaction (cs.HC), Robotics (cs.RO) |
| Authors | Parisa Ghanad Torshizi, Stacy Marsella |
| Submitted | August 24, 2026 |
| DOI | 10.48550/arXiv.2608.22731 |
Executive Summary
传统虚拟代理的非言语行为生成系统通常以话语作为输入,并生成用于强调或图解言语通道内容的非言语行为。然而,人类的非言语行为受到的影响远不止言语内容本身,它还受到说话者角色、人际关系、社会背景以及互动者认知与情绪状态的影响。因此,非言语通道可能会强化、削弱、限定甚至矛盾于言语通道。它还可能暴露出在言语中被隐藏或仅被间接暗示的内部状态,包括可能附带于即时互动的内心情绪“泄漏”。
Traditional nonverbal behavior generation systems for virtual agents focus heavily on alignment—translating an utterance into gestures and expressions that simply emphasize or illustrate the spoken words. However, human communication is far more complex; nonverbal actions are shaped by interpersonal relationships, speaker roles, social context, and internal emotional states.
Consequently, nonverbal cues can reinforce, weaken, qualify, or completely contradict verbal messaging. They may also reveal hidden or incidental internal states, such as emotional "leakage." To build more realistic, human-like virtual agents—especially for sensitive training scenarios like counseling simulations—researchers need models that capture this nuanced interplay.
建模这种更丰富的言语与非言语行为关系,对于设计表现出逼真、类人行为的虚拟代理至关重要。这在需要细致社会解释的训练语境中尤为关键,例如涉及虚拟病人的咨询模拟。借鉴埃克曼的言语-非言语关系框架,我们提出了一种分类体系,涵盖了言语与非言语行为之间可能出现不匹配的各类情况。随后,我们探讨了利用大语言模型实现这些行为的替代方法,重点研究大语言模型是否能够从给定的对话和社会互动语境中,选择出上下文恰当的不匹配言语和非言语行为。最后,我们在人类被试研究中对生成的行为进行了评估,以评估植入虚拟人的语境驱动型非言语行为是否对观察者产生了预期的效果。
Drawing upon Ekman's framework of verbal-nonverbal relationships, this paper proposes a taxonomy of behavioral mismatches. It investigates whether Large Language Models (LLMs) can effectively select contextually appropriate, incongruent verbal and nonverbal behaviors based on a dialogue interaction, and evaluates these embodied behaviors through human-subject studies.
Abstract
Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel may reinforce, weaken, qualify, or even contradict the verbal channel. It may also reveal internal states that are hidden or only indirectly implied in speech, including emotional "leakage" that may be incidental to the immediate interaction.
Modeling this richer relationship between verbal and nonverbal behavior is important for designing virtual agents that exhibit realistic, human-like behavior. It is especially critical in training contexts that require nuanced social interpretation, such as counseling simulations involving virtual patients. Drawing on Ekman's framework of verbal nonverbal relationships, we propose a taxonomy of categories in which mismatches between verbal and nonverbal behavior can occur. We then examine alternative approaches for realizing these behaviors using large language models, focusing on whether LLMs can select contextually appropriate mismatched verbal and nonverbal behaviors from a given dialogue and social interaction context. Finally, we evaluate the resulting behaviors in a human-subject study, assessing whether context-driven nonverbal behavior, when embodied in a virtual human, produces the intended effects on observers.
Key Highlights & Contributions
- 行为分类学: 建立了一个基于埃克曼心理学理论的结构化框架,对言语和非言语沟通渠道之间的不匹配和不一致情况进行分类。
- 大语言模型集成: 探讨了大语言模型解释对话语境并选择微妙且符合语境的不一致行为的能力,而不仅仅依赖于标准的直接强化。
- 人类被试评估: 评估观察者如何看待表现出真实、语境驱动的行为不一致的虚拟人,从而在诸如咨询模拟等应用场景中验证该方法的有效性。
- Behavioral Taxonomy: Establishes a structured framework—grounded in Ekman's psychological theories—categorizing mismatches and incongruencies between verbal and nonverbal communication channels.
- LLM Integration: Explores the capability of Large Language Models to interpret dialogue context and select nuanced, context-appropriate incongruent behaviors rather than relying on standard, direct reinforcement.
- Human-Subject Evaluation: Assesses how observers perceive virtual humans that exhibit realistic, context-driven behavioral incongruence, validating the approach in applied settings like counseling simulations.
Access Full-Text & Resources
- PDF 下载: 查看 PDF
- HTML 版本: arXiv HTML (实验性)
- 源码文件: TeX 源码
- 许可证: 知识共享署名 4.0

- PDF Download: View PDF
- HTML Version: arXiv HTML (Experimental)
- Source Files: TeX Source
- License: Creative Commons Attribution 4.0