跳转至

大语言模型在医学推理中展现出元认知敏感性

文章背景与核心概要

本研究探讨了大语言模型(LLM)在执行临床推理时,是否具备元认知敏感性——即将其置信度水平与证据质量及诊断不确定性进行对齐的能力。研究团队开发了一个受心理物理学启发、可控的临床基准测试,聚焦于阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关认知障碍(DRCI)的鉴别诊断,并通过45个合成病例分析了模型的诊断选择与置信度行为。

研究结果表明,模型表现出了一定程度的元认知敏感性(置信度随证据强度和信息完整性的变化而相应调整),但在中等难度且存在冲突的病例中也暴露出局部的校准失效(模型在倾向于某一诊断时保持了不合理的过高置信度)。这项研究为评估医学大语言模型的证据敏感性、元认知敏感性和局部校准失效建立了一个可复现的框架,并强调了直接测量置信度质量的重要性。


摘要

Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and \(\text{AUROC}_2\) was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.

大语言模型(LLMs)在医学领域的评估和应用日益广泛,但其临床实用性不仅取决于回答的准确性,还取决于其置信度是否能够准确反映证据质量和不确定性。我们开发了一个受心理物理学启发、可控的临床基准测试,用于检验医学大语言模型中的诊断选择与置信度行为。该基准测试聚焦于可能的阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关认知障碍(DRCI)之间的鉴别。我们生成了45个合成病例短文,其证据强度、冲突证据以及缺失信息各不相同。每个短文在三种提示词变体下进行测试,共产生135次试验。在使用 gpt-4.1-nano 进行的预实验中,所有试验均输出了有效的结构化结果。在强制选择试验中,诊断准确率为93.5%,平均置信度为78.4%,\(\text{AUROC}_2\) 为0.876。在调整了证据强度和提示词格式后,置信度随着证据与诊断边界距离的增加而增加,在信息缺失时下降,且正确试验中的置信度高于错误试验。这些发现表明模型具有局部的元认知敏感性,而非全局无参考价值的置信度。然而,错误集中在中等难度、存在冲突的 AT-NCD 病例中,在这些病例中,模型转向 DRCI 并保留了超出经验准确性所支持的过高置信度。模型比较表明,应直接测量置信度质量,而不应仅从基准准确性或模型能力推断。本研究为评估医学大语言模型中的证据敏感性、元认知敏感性和局部校准失效建立了一个可复现的框架。


元数据

  • arXiv ID: arXiv:2608.14552 [cs.AI]
  • DOI: 10.48550/arXiv.2608.14552
  • 作者: Ahmad Nazzal
  • 提交时间: 2026年5月4日
  • 学科分类: 人工智能 (cs.AI)

全文与访问链接


外部参考与引用