文章背景与核心概要
本文探讨了当输入相同的中心凹视网膜图像时,多模态大模型(MLLMs)在视觉搜索时的行为模式是否与人类相似。通过在目标导向搜索任务(COCO-Search18)中,以人类匹配的注视轨迹(scanpath)逐个注视点驱动三个通用 MLLM,作者从目标存在决策、搜索效率和视线过程三个维度评估了模型的表现。
研究发现,虽然 MLLM 在决策准确率和目标获取方面能够匹配甚至超越人类,但它们的注视过程与人类存在显著分歧。与人类不同,MLLM 表现出低熵、大振幅、自洽的注视轨迹,这表明它们采用的是单次通过(single-pass)、非串行的架构。这一发现表明,标准的答案对齐和显著性指标无法可靠地证明模型具备人类般的视觉,这意味着当前的零样本模型适用于结果和空间查询,但根本不适用于时间或过程层面的研究。
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
📌 Summary
📌 Summary
This paper investigates whether Multimodal Large Language Models (MLLMs) search visual scenes like humans do when provided with the same foveated input. By driving three general-purpose MLLMs fixation-by-fixation through human-matched scanpaths on goal-directed search tasks (COCO-Search18), the authors evaluate model performance across three axes: target presence decisions, search efficiency, and gaze process.
While MLLMs match or exceed humans in decision accuracy and target acquisition, their gaze processes diverge significantly. Unlike humans, MLLMs exhibit low-entropy, large-amplitude, self-consistent scanpaths indicative of a single-pass, non-serial architecture. The findings reveal that standard answer-alignment and saliency metrics cannot reliably certify human-like vision, making current zero-shot models suitable for outcome and spatial queries, but fundamentally unsuited for temporal, process-level investigations.
📋 Metadata
📋 Metadata
- arXiv ID: arXiv:2608.16514 [cs.CV]
- Subjects: Computer Vision and Pattern Recognition (
cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multimedia (cs.MM)- Submission Date: August 17, 2026
- Conference Status: Accepted at the 3rd HCV workshop at ECCV 2026 (12 pages main text, 16 pages supplementary material)
- DOI: 10.48550/arXiv.2608.16514
👥 Authors
👥 Authors
- Mohamed Amine Kerkouri
- Marouane Tliba
- Aladine Chetouani
- Ulas Bagci
- Alessandro Bruno
📄 Abstract
📄 Abstract
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself.
The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.
🔗 Full-Text & Resources
🔗 Full-Text & Resources
- PDF: View PDF
- HTML (Experimental): arXiv HTML View
- TeX Source: arXiv Source File
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0