大语言模型时代的医学推理:增强技术与应用系统的全面综述
文章背景与核心概要
本篇系统性综述深入探讨了大语言模型(LLMs)在临床医学中的应用,重点填补了“单步答案生成”与“透明、可验证的医学推理”之间存在的关键技术鸿沟。文章全面梳理并总结了2022年至2025年间的60项开创性研究,提出了一套涵盖训练期策略与测试期机制的推理增强技术分类法。
此外,该研究还评估了这些技术在多模态数据类型、核心临床领域以及不断演进的评估框架中的部署情况,并指出了未来亟待解决的挑战,包括忠实度与合理性之间的差距(faithfulness-plausibility gap)以及原生多模态等前沿方向。这项工作为构建高效、稳健且符合社会技术责任的医学人工智能指明了发展路径。
📋 总结
This systematic review investigates the application of Large Language Models (LLMs) in clinical medicine, addressing the critical gap between single-step answer generation and transparent, verifiable medical reasoning. Spanning 60 seminal studies from 2022 to 2025, the paper proposes a comprehensive taxonomy of reasoning enhancement techniques—dividing them into training-time strategies and test-time mechanisms. It further evaluates their deployment across multi-modal data types, core clinical domains, and evolving evaluation frameworks, while outlining future challenges such as the faithfulness-plausibility gap and native multimodality.
📌 文档元数据
- arXiv ID: arXiv:2508.00669 [cs.CL]
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)- Submission Timeline:
- Submitted on August 1, 2025 (v1)
- Last revised on September 3, 2026 (v2)
- DOI: 10.48550/arXiv.2508.00669
👥 作者列表
- Zizhan Ma
- Wenxuan Wang
- Meidan Ding
- Shiyi Zheng
- Shengyuan Liu
- Jie Liu
- Jiaming Ji
- Linlin Shen
- Yixuan Yuan
- Wenting Chen
📖 摘要
大语言模型(LLMs)在医学领域的普及展现出令人瞩目的能力,然而在执行系统性、透明且可验证的推理方面(这是临床实践的基石),依然存在巨大的技术鸿沟。这促使业界从单纯的单步答案生成,转向开发专门针对医学推理设计的大语言模型。
本文对这一新兴领域进行了首次系统性综述,突出了以下核心贡献: 1. 增强技术分类法: 将其划分为训练期策略(如监督微调、强化学习)和测试期机制(如提示工程、多智能体系统)。 2. 多模态与临床应用: 分析了跨数据模态(文本、图像、代码)及关键临床工作流(诊断、教育和治疗规划)的集成应用。 3. 评估基准: 追踪了从简单准确率指标向复杂的推理质量和视觉可解释性评估体系的演变过程。 4. 未来展望: 识别了如忠实度与合理性差距等关键障碍,指明了通往高效、稳健和符合社会技术责任的医学AI的演进路径。
The proliferation of Large Language Models (LLMs) in medicine has enabled impressive capabilities, yet a critical gap remains in their ability to perform systematic, transparent, and verifiable reasoning, a cornerstone of clinical practice. This has catalyzed a shift from single-step answer generation to the development of LLMs explicitly designed for medical reasoning.
This paper provides the first systematic review of this emerging field, highlighting key contributions: 1. Taxonomy of Enhancement Techniques: Categorized into training-time strategies (e.g., supervised fine-tuning, reinforcement learning) and test-time mechanisms (e.g., prompt engineering, multi-agent systems). 2. Multimodal and Clinical Applications: Analyzing integration across data modalities (text, image, code) and vital clinical workflows (diagnosis, education, and treatment planning). 3. Evaluation Benchmarks: Tracing the evolution from simple accuracy metrics to sophisticated assessments of reasoning quality and visual interpretability. 4. Future Outlook: Identifying critical obstacles like the faithfulness-plausibility gap, highlighting paths toward efficient, robust, and sociotechnically responsible medical AI.