波兰语医学视觉问答:视觉语言模型未能充分利用视觉证据
文章背景与核心概要
随着多模态人工智能的迅猛发展,视觉语言模型(VLMs)在医学领域的应用受到了广泛关注。然而,现有的大多数医学多模态基准测试主要集中在英语环境,且对模型在复杂临床推理中实际利用视觉证据的程度缺乏深入剖析。本文引入了一个全新的波兰语医学视觉问答(VQA)基准测试,该基准直接取材于面向执业医师和牙医专科认证的波兰国家考试真实真题。
研究团队通过对多款开源及商业视觉语言模型的全面评估发现,当前最先进的模型在处理此类高质量医学多模态任务时仍面临严峻挑战。更引人深思的是,消融实验表明,模型严重过度依赖文本问题而非图像中的视觉证据,在以图像为主导的问题上表现最差。此外,模型仅凭候选答案选项就能获得高於随机猜测的准确率,揭示了其在面对缺失关键任务组件时依然存在的“捷径学习(Shortcut Learning)”现象。
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
波兰语医学视觉问答:视觉语言模型未能充分利用视觉证据
Authors: Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
作者: Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
Primary Subject: Artificial Intelligence (cs.AI)
主要学科: 人工智能 (cs.AI)
arXiv ID: arXiv:2608.12928
提交时间: 2026年8月13日
Executive Summary
执行摘要
This paper introduces a novel benchmark for Polish-language Medical Visual Question Answering (VQA), compiled from authentic Polish Board Certification Examination questions designed for licensed physicians and dentists.
本文引入了一个全新的波兰语医学视觉问答(VQA)基准,该基准汇编自专为持证医师和牙医设计的波兰专科医师认证考试的真实考题。
Key Findings & Insights:
核心发现与洞察:
- Current Limitations: Vision-Language Models (VLMs) struggle significantly with this multimodal task. The highest-performing model achieved 79.0% accuracy on the full VQA set. Only one model (GPT-5.6) outperformed the approximate human baseline on subsets with candidate responses; all other evaluated models fell short of human performance.
- 当前局限性: 视觉语言模型(VLMs)在此类多模态任务中表现出显著的困难。表现最好的模型在完整的VQA数据集上取得了 79.0% 的准确率。在包含候选答案的子集上,仅有一个模型(GPT-5.6)超过了大致的人类基线水平;所有其他接受评估的模型均未达到人类的表现水平。
- Visual Underutilization: Through ablation tests (omitting images, questions, or candidate answers), the researchers discovered that models rely much more heavily on the text of the question than on visual evidence. They also performed worst on image-dominant questions.
- 视觉证据利用不足: 通过消融实验(省略图像、问题或候选答案),研究人员发现模型对问题文本的依赖远大于对视觉证据的依赖。它们在以图像为主导的问题上的表现也最差。
- Shortcut Learning: Across both text-only QA and multimodal VQA setups, models achieved above-chance accuracy purely from the answer choices alone, revealing that models can exploit superficial shortcuts even when critical task components are missing.
- 捷径学习: 无论是纯文本问答(QA)还是多模态VQA设置,模型仅凭答案选项本身就能获得高於随机猜测的准确率,这表明即使缺少关键的任务组件,模型也能利用表面的捷径。
Abstract
摘要
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves \(79.0\%\) accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
我们引入了一个波兰语医学视觉问答(VQA)基准,该基准构建自面向寻求专科认证的持证医师和牙医的波兰专科医师认证考试试题。该基准包含跨越多个医学专业和视觉领域的含图问题,以及一个纯文本问答(QA)对照集。我们评估了面向波兰语的、通用开源的以及商业的视觉语言模型。这项任务仍然充满挑战:表现最好的模型在完整的VQA数据集上达到了 \(79.0\%\) 的准确率,并且只有 GPT-5.6 在提供候选答案的子集上超过了人类的大致参考水平;所有其他评估模型的表现均逊于人类。为了评估视觉扎根(visual grounding),我们对比了完整输入与省略图像、问题或二者的配置,并按图像重要性对问题进行了分类。模型从问题文本中获取的有用信息多于从图像中获取的信息,并且在图像主导的问题上表现更差。尽管如此,在QA和VQA任务中,它们仅凭答案选项本身就能获得高於随机猜测的准确率,这表明即使缺少关键的任务组件,非微不足道的性能依然可以保持。
Links and Resources
链接与资源
- Full-Text Access:
- 全文访问:
- View PDF
- 查看 PDF
- HTML Version (Experimental)
- HTML 版本(实验性)
- TeX Source
- TeX 源码
- License: Creative Commons Attribution 4.0

- 许可证: 知识共享署名 4.0

- Citations & References:
- 引用与参考:
- NASA ADS
- 谷歌学术 (Google Scholar)
- 语义学者 (Semantic Scholar)