文章背景与核心概要
本文探讨了临床视觉-语言模型(VLM)在处理包含神经影像背景提示时的可靠性。作者评估了 12 个开源 VLM 在两个临床神经影像队列(针对情感障碍和认知衰退)上的表现,发现仅在提示中加入神经影像引用(甚至在提供任何图像之前),就能让较小模型的性能指标显著提升(F1 值最高提升 0.66),同时校准度也有所改善。然而,专家个案研究表明,模型的真实性(Faithfulness)仍然很差,经常引入未经证实的临床细节。此外,单模型偏好对齐尝试抑制了引用 MRI 的行为,但最终损害了性能优势,从而未能解决表面多模态融合的核心问题。
Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions
Summary
This paper investigates the reliability of clinical vision-language models (VLMs) when processing prompts that mention neuroimaging context. Evaluating 12 open-weight VLMs on two clinical neuroimaging cohorts (for affective disorders and cognitive decline), the authors discover that simply adding a neuroimaging reference to the prompt—even before any image is provided—causes smaller models to significantly boost their performance metrics (gaining up to 0.66 F1) and calibration. However, expert case studies reveal that model faithfulness remains poor, often introducing unverified clinical details. Furthermore, single-model preference alignment attempts suppress MRI-referencing behavior but ultimately compromise the performance advantage, leaving the core issue of superficial multimodal integration unresolved.
Metadata
- arXiv ID: arXiv:2603.28387 [cs.AI]
- DOI: 10.48550/arXiv.2603.28387
- Primary Subject: Artificial Intelligence (
cs.AI) - Secondary Subject: Machine Learning (
cs.LG) - Conference Acceptance: Accepted to EMNLP 2026 Main Conference
- Authors: Doan Nam Long Vu, Simone Balloccu
- Submission History:
- [v1] Mon, 30 Mar 2026
- [v2] Thu, 18 Jun 2026
- [v3] Fri, 28 Aug 2026 (Current)
Abstract
Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitive decline. Both cohorts include structural magnetic resonance imaging (MRI) acquired under their original research protocols. Prior work does not establish the included neuroimaging inputs as reliable stand-alone diagnostic evidence for the present tasks.
Nevertheless, when neuroimaging context is introduced, smaller VLMs gain up to 0.66 F1 under the evaluated augmented conditions, becoming competitive with models an order of magnitude larger. Confidence estimation shows that most of the calibration improvement for the analyzed smaller models occurs after the MRI reference is added to the prompt, before any image is supplied. Our preliminary expert case study finds that faithfulness remains low in every condition examined, with the reviewed model introducing unverified clinical details. Finally, in our single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved. These results caution against reading surface metric gains as evidence of true multimodal integration, with direct implications for clinical VLM deployment.
值得信赖的临床人工智能必须使用真实的证据,避免依赖表面层次的伪影。我们在两个临床神经影像队列上评估了 12 个开源视觉-语言模型(VLM),用于情感障碍和认知衰退的二元分类。这两个队列都包含在其原始研究协议下获取的结构磁共振成像(MRI)。先前的工作并未确定所包含的神经影像输入可作为当前任务的可靠独立诊断证据。
然而,当引入神经影像背景时,较小的 VLM 在评估的增强条件下获得了高达 0.66 的 F1 提升,从而能够与大一个数量级的模型相竞争。置信度估计表明,所分析的较小模型的大部分校准改进都发生在将 MRI 引用添加到提示中之后、提供任何图像之前。我们的初步专家个案研究发现,在检查的每种条件下,模型的真实性仍然很低,审阅的模型引入了未经证实的临床细节。最后,在我们的单模型干预中,偏好对齐抑制了引用 MRI 的行为,但降低了增强条件下的优势,从而使潜在问题未能解决。这些结果警告人们,不要将表面指标的提升解读为真正的多模态融合的证据,这对临床 VLM 的部署具有直接的影响。
Associated Links & Resources
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution 4.0 International

- External Citations:
- NASA ADS
- Google Scholar
- Semantic Scholar