文章背景与核心概要
本文探讨了多模态大语言模型(MLLMs)在处理违背常识预期的视觉场景时所表现出的局限性。尽管 MLLMs 在主流视觉任务中表现出色,但作者证明它们经常遭受“语言偏见”的困扰,即模型往往优先选择统计学上常见的文本描述,而不是实际的视觉证据。为了量化这一现象,研究人员推出了 CAIT 基准测试,其中包含 400 个具有反直觉动作的高保真合成场景(例如,“兔子追赶老虎”)。
研究发现,虽然人类在 CAIT 基准测试上的准确率约为 95%,闭源商业模型(如 Claude 和 Gemini)最高可达 88%,但标准的开源指令微调模型仅表现出随机猜测的水平。开源模型倾向于用常识性的文本先验来覆盖异常的视觉信号。该研究为多模态模型在视觉真实性与世界知识之间的对齐问题提供了重要的分析与改进方向。
眼见为实还是先入为主:评估开源多模态大语言模型在反直觉场景中的语言偏见 (Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes)
作者: Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv: 2601.07737 [cs.CV]
提交时间: 2026年1月12日 (v1), 2026年8月25日 (v3)
摘要 (Summary)
本文探讨了多模态大语言模型(MLLMs)在处理违背常识预期的视觉场景时所表现出的局限性。尽管 MLLMs 在主流视觉任务中表现出色,但作者证明它们经常遭受“语言偏见”的困扰,即模型往往优先选择统计学上常见的文本描述,而不是实际的视觉证据。为了量化这一现象,研究人员推出了 CAIT 基准测试,其中包含 400 个具有反直觉动作的高保真合成场景(例如,“兔子追赶老虎”)。
This paper investigates the limitations of Multimodal Large Language Models (MLLMs) when processing visual scenes that contradict common-sense expectations. While MLLMs excel at mainstream visual tasks, the authors demonstrate that they often suffer from a "language bias," where the model prioritizes statistically common text descriptions over actual visual evidence. To quantify this, the researchers introduced CAIT, a benchmark of 400 high-fidelity synthetic scenes featuring counter-intuitive actions (e.g., "a rabbit chasing a tiger").
核心发现 (Key Findings)
- 性能差距: 人类在 CAIT 基准测试上的准确率约为 95%,闭源商业模型(如 Claude 和 Gemini)最高可达到 88%,而标准的开源指令微调模型则仅表现出随机猜测的水平。
- “语言先验”问题: 开源模型倾向于用常识性的文本先验来覆盖异常的视觉信号。
- 思维链(CoT)的局限性: 尽管思维链推理可以提高准确率,但它引入了一种新的失效模式——“过度思考”,即模型仅仅因为违反了现实世界的物理规律而拒绝有效的视觉证据。
- 缓解策略: 研究得出结论,针对性的微调和结构化提示可以有效减少对语言先验的依赖,使模型能够更好地将推理建立在视觉现实之上。
- Performance Gap: While humans achieve ~95% accuracy and proprietary models (like Claude and Gemini) reach up to 88% on the CAIT benchmark, standard open-source instruction-tuned models perform at chance levels.
- The "Language Prior" Problem: Open-source models tend to override anomalous visual signals with common-sense text priors.
- Chain-of-Thought (CoT) Limitations: While CoT reasoning can improve accuracy, it introduces a new failure mode: "overthinking," where models reject valid visual evidence simply because it violates real-world physical laws.
- Mitigation Strategies: The study concludes that targeted fine-tuning and structured prompting can effectively reduce reliance on language priors, allowing models to better ground their reasoning in visual reality.
访问与资源 (Access & Resources)
引用与元数据 (Citation & Metadata)
- 主要主题: 计算机视觉与模式识别 (cs.CV)
- 次要主题: 人工智能 (cs.AI)
- DOI: 10.48550/arXiv.2601.07737
- Primary Subject: Computer Vision and Pattern Recognition (cs.CV)
- Secondary Subject: Artificial Intelligence (cs.AI)
- DOI: 10.48550/arXiv.2601.07737