跳转至

文章背景与核心概要

本文深入探讨了多模态大语言模型(MLLMs)的内部机制,旨在理解它们如何编码和定位不同的视觉概念。通过提出一种基于激活干涉(activation steering)的创新因果框架,作者对四大类视觉概念进行了系统性的干涉实验。

核心研究发现包括:实体知识表现出明显的局域化特征,而抽象概念则在整个网络中呈全局分布式编码;模型深度的增加对于编码复杂的抽象概念至关重要,而实体知识则在不同模型规模下均保持高度局域化;通过逆向干涉发现,阻断显式输出会导致潜在激活激增,这表明感知和生成之间存在补偿机制;此外,尽管MLLM能够识别几何关系,但它们仅将其视为静态视觉特征,并未触发问题解决所需的程序化执行。


Causal Probing for Internal Visual Representations in Multimodal Large Language Models

arXiv: 2605.05593 [cs.AI]
Authors: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
Submission History: Submitted on 7 May 2026; Last revised 1 September 2026 (v2).

arXiv: 2605.05593 [cs.AI]
Authors: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
Submission History: Submitted on 7 May 2026; Last revised 1 September 2026 (v2).


Summary

This paper investigates the internal mechanisms of Multimodal Large Language Models (MLLMs) to understand how they encode and ground distinct visual concepts. Using a novel causal framework based on activation steering, the authors systematically intervene across four visual concept categories.

Key findings include: * Concept Encoding Divergence: Entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. * Driver of Scaling Laws: Increasing model depth is essential for encoding complex abstract concepts, whereas entities maintain a high degree of localization regardless. * Perception vs. Generation: Reverse steering shows that blocking explicit output triggers a surge in latent activations, pointing to a compensatory mechanism between perception and generation. * Visual Reasoning Disconnect: While MLLMs recognize geometric relations, they treat them merely as static visual features rather than triggering the procedural execution necessary for problem-solving.

Summary

This paper investigates the internal mechanisms of Multimodal Large Language Models (MLLMs) to understand how they encode and ground distinct visual concepts. Using a novel causal framework based on activation steering, the authors systematically intervene across four visual concept categories.

Key findings include: * Concept Encoding Divergence: Entity knowledge is distinctively localized, whereas abstract concepts are globally distributed across the network. * Driver of Scaling Laws: Increasing model depth is essential for encoding complex abstract concepts, whereas entities maintain a high degree of localization regardless. * Perception vs. Generation: Reverse steering shows that blocking explicit output triggers a surge in latent activations, pointing to a compensatory mechanism between perception and generation. * Visual Reasoning Disconnect: While MLLMs recognize geometric relations, they treat them merely as static visual features rather than triggering the procedural execution necessary for problem-solving.


Metadata & References

Metadata & References