文章背景与核心概要
尽管多模态大语言模型(MLLMs)在各项任务中取得了巨大成功,但其编码和基础化(grounding)不同视觉概念的内部运作机制依然知之甚少。本文引入了一种基于激活引导的因果框架,用于主动探测和操纵内部视觉表示。通过对四个视觉概念类别进行系统性干预,该研究揭示了关于 MLLM 架构和推理机制的关键洞察。
研究发现,实体知识具有明显的局部化特征,而抽象概念则在整个网络中呈全局分布式存储;这种编码差异解释了模型规模定律(scaling laws)的一个底层机制:增加模型深度对于编码复杂的分布式抽象概念至关重要,而实体定位则保持稳定。此外,逆向引导实验揭示了感知与生成之间的补偿机制:阻断显式输出会触发潜在激活的激增。最后,研究暴露了感知与推理之间的脱节:尽管 MLLMs 能够识别几何关系,但它们仅将其作为静态视觉特征处理,而未能触发解决复杂问题所需的程序化执行。
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
arXiv: 2605.05593 [cs.AI]
Accepted at: EMNLP 2026 Main
Submission History: Submitted on May 7, 2026; last revised September 3, 2026 (v3).
Authors: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
📋 Summary
Despite the widespread success of Multimodal Large Language Models (MLLMs), the internal mechanisms governing how they encode and ground visual concepts remain poorly understood. This paper introduces a causal framework based on activation steering to actively probe and manipulate internal visual representations.
Through systematic interventions across four visual concept categories, the study reveals critical insights into MLLM architecture and reasoning: 1. Concept Encoding Divergence: Entity knowledge is distinctly localized, whereas abstract concepts are globally distributed across the network. 2. Driver of Scaling Laws: Increasing model depth is essential for encoding distributed abstract concepts, while entities maintain stable localization. 3. Perception-Generation Compensatory Mechanism: Reverse steering demonstrates that blocking explicit output triggers a surge in latent activations. 4. Perception-Reasoning Disconnect: While MLLMs can recognize geometric relations, they treat them as static visual features rather than triggering the necessary procedural execution for complex problem-solving.
📝 Abstract
Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode and ground distinct visual concepts remain poorly understood. To unravel these mechanisms, we propose a causal framework based on activation steering to actively probe and manipulate internal visual representations. Through systematic intervention across four visual concept categories, our results reveal a divergence in concept encoding: entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. Critically, this divergence uncovers a mechanistic driver of scaling laws: increasing model depth is indispensable for encoding distributed and complex abstract concepts, whereas entities maintain a consistently high degree of localization. Furthermore, reverse steering uncovers that blocking explicit output triggers a surge in latent activations, exposing a compensatory mechanism between perception and generation. Finally, by extending our analysis to visual reasoning, we expose a disconnect between perception and reasoning: although MLLMs successfully recognize geometric relations, they treat them merely as static visual features, failing to trigger the procedural execution necessary for solving problems.
🔗 Quick Links
- Full-Text Options: View PDF | HTML (Experimental) | TeX Source
- External Citations: Google Scholar | Semantic Scholar | NASA ADS
- Associated Code & Tools: Hugging Face | CatalyzeX Code Finder | alphaXiv