听觉、视觉与追踪:全模态语言模型的时空视听声学事件推理
文章背景与核心概要
在多媒体环境中理解动态声源,需要同时确定是什么产生了声音、声源位于何处以及它随时间如何移动。然而,现有的音频语言模型往往将声音片段视为全局声学事件,而视觉语言模型则缺乏定位和追踪单个声源所需的空间音频线索。
为了填补这一空白,本文作者推出了 ST-OmniQA,这是一个包含 40K 全景视频和 400K 问答对的大规模时空视听问答基准。在此基准的基础上,他们进一步提出了 ST-Omni-R1,这是一个通过渐进式课程学习和推理树强化学习训练而成的全模态语言模型,能够有效整合空间音频与视觉上下文。
实验结果表明,ST-Omni-R1 在四个能力级别上取得了 77.83% 的平均语义准确率(作为对比,表现最好的基线模型仅为 37.28%),展现出了卓越的时空视听推理与追踪能力。
📌 摘要
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources.
理解动态声源需要同时确定是什么产生了声音、声源位于何处以及它随时间如何移动。然而,现有的音频语言模型往往将音频片段表示为全局声学事件,而视觉语言模型则缺乏定位和追踪单个声源所需的空间音频线索。
To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering: * Sound-event recognition * Direction of arrival * Source distance * Motion trajectories * Temporally grounded audio-visual reasoning
为了评估这一缺失的能力,我们推出了 ST-OmniQA,这是一个时空视听问答基准,由全景视频以及与之同步的、包含移动声源的一阶高保真立体声(FOA)音频构建而成。它包含 40K 个视频和 400K 个问答对,按四个能力级别进行组织,涵盖: * 声学事件识别 * 到达方向 * 声源距离 * 运动轨迹 * 基于时间落脚点的视听推理
Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning.
在此基准的基础上,我们提出了 ST-Omni-R1,它将 FOA 派生的语义和轨迹表征与全景视觉上下文相结合,并通过渐进式课程学习和推理树强化学习进行训练。
Performance: ST-Omni-R1 achieves 77.83% average semantic accuracy across the four levels, compared with 37.28% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.
性能表现: ST-Omni-R1 在四个级别上取得了 77.83% 的平均语义准确率,相比之下,表现最好的评估基线模型为 37.28%。在三个公共空间音频基准上的结果进一步表明,其学到的空间和运动表征能够泛化并迁移到 ST-OmniQA 之外的场景。
📑 论文详情与元数据
- Primary Subject: Artificial Intelligence (
cs.AI) - Cite as:
arXiv:2608.09435[cs.AI] - Full-Text Links:
- View PDF
- HTML Version (Experimental)
- TeX Source
- 主要学科: 人工智能 (
cs.AI)- 引用格式:
arXiv:2608.09435[cs.AI]- 全文链接:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
🔗 外部资源与引用
- Code, Data & Media: Available via integrations with Hugging Face, CatalyzeX, and DagsHub (accessible through the arXivLabs interface).
- Reference Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
- 代码、数据与媒体: 可通过与 Hugging Face、CatalyzeX 和 DagsHub 的集成获取(通过 arXivLabs 界面访问)。
- 参考工具:
- Google Scholar
- Semantic Scholar
- NASA ADS
