文章背景与核心概要
本文深入探讨了多模态大语言模型(MLLMs)的内部机制,旨在理解它们如何编码和定位不同的视觉概念。通过提出一种基于激活干涉(activation steering)的创新因果框架,作者对四大类视觉概念进行了系统性的干涉实验。
核心研究发现包括:实体知识表现出明显的局域化特征,而抽象概念则在整个网络中呈全局分布式编码;模型深度的增加对于编码复杂的抽象概念至关重要,而实体知识则在不同模型规模下均保持高度局域化;通过逆向干涉发现,阻断显式输出会导致潜在激活激增,这表明感知和生成之间存在补偿机制;此外,尽管MLLM能够识别几何关系,但它们仅将其视为静态视觉特征,并未触发问题解决所需的程序化执行。
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
arXiv: 2605.05593 [cs.AI]
Authors: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
Submission History: Submitted on 7 May 2026; Last revised 1 September 2026 (v2).
arXiv: 2605.05593 [cs.AI]
Authors: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
Submission History: Submitted on 7 May 2026; Last revised 1 September 2026 (v2).
Summary
This paper investigates the internal mechanisms of Multimodal Large Language Models (MLLMs) to understand how they encode and ground distinct visual concepts. Using a novel causal framework based on activation steering, the authors systematically intervene across four visual concept categories.
Key findings include: * Concept Encoding Divergence: Entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. * Driver of Scaling Laws: Increasing model depth is essential for encoding complex abstract concepts, whereas entities maintain a high degree of localization regardless. * Perception vs. Generation: Reverse steering shows that blocking explicit output triggers a surge in latent activations, pointing to a compensatory mechanism between perception and generation. * Visual Reasoning Disconnect: While MLLMs recognize geometric relations, they treat them merely as static visual features rather than triggering the procedural execution necessary for problem-solving.
Summary
This paper investigates the internal mechanisms of Multimodal Large Language Models (MLLMs) to understand how they encode and ground distinct visual concepts. Using a novel causal framework based on activation steering, the authors systematically intervene across four visual concept categories.
Key findings include: * Concept Encoding Divergence: Entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. * Driver of Scaling Laws: Increasing model depth is essential for encoding complex abstract concepts, whereas entities maintain a high degree of localization regardless. * Perception vs. Generation: Reverse steering shows that blocking explicit output triggers a surge in latent activations, pointing to a compensatory mechanism between perception and generation. * Visual Reasoning Disconnect: While MLLMs recognize geometric relations, they treat them merely as static visual features rather than triggering the procedural execution necessary for problem-solving.
Metadata & References
- Primary Subject: Artificial Intelligence (
cs.AI) - DOI: 10.48550/arXiv.2605.05593
- Full-Text Links: View PDF | HTML (Experimental) | TeX Source
- External Resources:
- Google Scholar
- Semantic Scholar
- NASA ADS
Metadata & References
- Primary Subject: Artificial Intelligence (
cs.AI)- DOI: 10.48550/arXiv.2605.05593
- Full-Text Links: View PDF | HTML (Experimental) | TeX Source
- External Resources:
- Google Scholar
- Semantic Scholar
- NASA ADS