文章背景与核心概要
长视频理解不仅需要检索孤立的事件,更需要追踪不断演变的故事情节并解读其中的隐性社会含义。然而,目前的视频基准测试很少将这些能力结合起来进行评估,特别是在高语境、非英语的媒体内容方面。
为了填补这一空白,作者推出了 NARU,这是一个专门的基准测试,旨在利用日本长视频评估叙事演进(Narrative evolution)与文化理解推理(Reasoning on cultural Understanding)。该基准包含来自 155 个视频的总计 1,46.8 小时的 1,481 个问题,涵盖四个叙事维度和五个文化维度。对多种模型配置的评估表明,当前的多模态大模型(MLLMs)在长距离叙事整合和基于文化的推理方面均存在显着的局限性。
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
arXiv: arXiv:2608.13210 [cs.CV]
Submitted: August 13, 2026
Authors: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
arXiv: arXiv:2608.13210 [cs.CV]
Submitted: August 13, 2026
Authors: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
📌 Summary
📌 Summary
长视频理解不仅需要检索孤立的事件;它还涉及追踪不断演变的故事情节和解释隐含的社会含义。然而,当前的视频基准很少共同评估这些能力,特别是对于高语境、非英语的媒体。
Long-form video understanding requires more than just retrieving isolated events; it involves tracking evolving storylines and interpreting implicit social meanings. However, current video benchmarks rarely evaluate these abilities jointly, especially for high-context, non-English media.
为了弥补这一差距,作者推出了 NARU,这是一个专门的基准,旨在利用日本长视频评估叙事演进和文化理解推理。该基准包含来自 155 个视频的 1,481 个问题,总时长达 146.8 小时,涵盖四个叙事维度和五个文化维度。对多个模型配置的评估揭示了当前多模态大模型(MLLMs)在长距离叙事整合和文化基础推理方面的重大局限性。
To bridge this gap, the authors introduce NARU, a specialized benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding using Japanese long-form videos. The benchmark comprises 1,481 questions across 155 videos totaling 146.8 hours, covering four narrative dimensions and five cultural dimensions. Evaluations across multiple model configurations reveal significant limitations in both long-range narrative integration and culturally grounded reasoning among current Multimodal Large Language Models (MLLMs).
👥 Authors & Contributions
👥 Authors & Contributions
- Yuheng Huang 与 Jianlang Chen (同等贡献)
- Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
- Yuheng Huang & Jianlang Chen (Contributed equally)
- Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
📊 Benchmark Overview
📊 Benchmark Overview
- 规模: 建立在 155 个视频基础上的 1,481 个问题。
- 总时长: 146.8 小时的极长内容。
- 范围:
- 4 个叙事维度
- 5 个文化维度
- 标注流程: 采用基于分层记忆的标注流程,将原始视频转换为结构化事件、叙事和文化标注,并结合任务导向的综合与迭代捷径移除。
- 验证: 通过涉及 68 名标注人员的两个母语者验证阶段进行验证。
- Scale: 1,481 questions grounded in 155 videos.
- Total Duration: 146.8 hours of extreme long-form content.
- Scope:
- 4 Narrative Dimensions
- 5 Cultural Dimensions
- Annotation Pipeline: Features a hierarchical memory-based annotation pipeline transforming raw video into structured events, narratives, and cultural annotations, combined with task-oriented synthesis and iterative shortcut removal.
- Verification: Validated through two native-speaker verification stages involving 68 annotators.
🔗 Links & Resources
🔗 Links & Resources
- 论文与 PDF: 在 arXiv 上查看 PDF | arXiv 摘要
- 项目网站: MA Labo NARU | InfiniMind 新闻发布
- 引用: Google Scholar | Semantic Scholar | NASA ADS
- Paper & PDF: View PDF on arXiv | arXiv Abstract
- Project Websites: MA Labo NARU | InfiniMind News Release
- Citations: Google Scholar | Semantic Scholar | NASA ADS