身份感知的人机交互动作描述生成
文章背景与核心概要
现有的“人机交互(HOI)动作描述生成系统”通常使用诸如“一个人”或“某人”之类的通用词汇来描述动作,未能将描述与特定的主体身份(Identity)相联系。为了弥补这一空白,本文引入了身份感知的人机交互动作描述生成(Identity-Aware Human-Object Interaction Motion Captioning)任务。该任务要求生成的描述能够明确指定独特的主体身份以及精确的HOI动作(例如,用 "Sub_ID lifts the chair" 代替 "A person lifts the chair")。
为了支持这一新任务,作者基于 BEHAVE 和 InterCap 数据集构建了全新的身份感知HOI动作描述,并提出了 ID-HOINet 网络。该创新网络能够从多视角的视频中进行学习,并通过两大核心组件执行单视角的身份感知HOI动作描述生成:1. 多视角身份-动作学习模块(MVIML):对时间阶段和相机视角之间的依赖关系进行建模,以捕获独特的身份和交互动作特征;2. 两阶段描述重写策略(TSCR):在推理过程中检索主体身份并生成与身份无关的HOI动作描述,随后利用预测出的身份对其进行重写,从而输出精确的结果。实验证明,ID-HOINet在该基准测试上取得了最先进(SOTA)的性能。
摘要 (Summary)
现有人机交互(HOI)动作描述生成系统通常使用“一个人”或“某人”等通用术语来描述动作,未能将描述落实到特定的主体身份上。为了弥补这一差距,本文引入了身份感知的人机交互动作描述生成任务,该任务要求生成的描述明确指定独特的主体身份和精确的HOI动作(例如,用 "Sub_ID lifts the chair" 代替 "A person lifts the chair")。
Existing human-object interaction (HOI) motion captioning systems typically describe actions using generic terms such as "a person" or "someone," failing to ground the descriptions in specific subject identities. To bridge this gap, this paper introduces the Identity-Aware Human-Object Interaction Motion Captioning task, which requires generated captions to explicitly specify both the unique subject identity and the precise HOI motion (e.g.,
"Sub_ID lifts the chair"instead of"A person lifts the chair").
为了支持这一任务,作者利用 BEHAVE 和 InterCap 数据集构建了新的身份感知HOI动作描述,并提出了 ID-HOINet。这个新颖的网络从多视角视频中学习,并通过两个主要组件执行单视角身份感知HOI动作描述生成: 1. 多视角身份-动作学习模块(MVIML):对时间阶段和相机视角的依赖关系进行建模,以捕获不同的身份和交互动作特征。 2. 两阶段描述重写策略(TSCR):在推理过程中检索主体身份并生成与身份无关的HOI动作描述,随后用预测的身份重写它们以获得精确的输出。
To support this task, the authors construct new identity-aware HOI motion captions using the BEHAVE and InterCap datasets and propose ID-HOINet. This novel network learns from multi-view videos and performs single-view identity-aware HOI motion caption generation using two main components: 1. Multi-View Identity-Motion Learning Module (MVIML): Models dependencies across temporal stages and camera viewpoints to capture distinct identity and interaction motion features. 2. Two-Stage Caption Rewriting Strategy (TSCR): Retrieves the subject identity and generates identity-agnostic HOI motion captions during inference, subsequently rewriting them with the predicted identity for precise output.
实验表明,ID-HOINet在该基准测试中取得了最先进的性能。
Experiments demonstrate that ID-HOINet achieves state-of-the-art performance on this benchmark.
元数据与出版详情 (Metadata & Publication Details)
- arXiv ID: arXiv:2608.20690 [cs.CV]
- 主要主题: 计算机视觉与模式识别 (
cs.CV) - 次要主题: 人工智能 (
cs.AI) - 提交时间: 2026年8月21日
- 作者:
- Yiming Wang
- Yonghao Dang
- Huilai Li
- Jiawei Tu
- Jianqin Yin
- 论文页数: 9页,3幅图
- arXiv ID: arXiv:2608.20690 [cs.CV]
- Primary Subject: Computer Vision and Pattern Recognition (
cs.CV)- Secondary Subjects: Artificial Intelligence (
cs.AI)- Submitted On: August 21, 2026
- Authors:
- Yiming Wang
- Yonghao Dang
- Huilai Li
- Jiawei Tu
- Jianqin Yin
- Paper Length: 9 pages, 3 figures