文章背景与核心概要
大语言模型(LLM)在医学诊断领域展现出了巨大的应用前景,但现有的提示词工程和多智能体系统通常局限于孤立的单次推理,缺乏对可复用临床经验的积累。为了填补这一空白,本文提出了 MACD(Multi-Agent Clinical Diagnosis) 框架,通过结构化的多智能体流水线让大模型实现临床知识的自学习——像人类医师一样总结、提炼并应用诊疗经验。
此外,作者还提出了一套 MACD-human 协同工作流,将大模型迭代会诊、裁判智能体以及人类监督有机结合。在全新构建的 MIMIC-MACD 队列(包含涵盖七种疾病的 4,390 个真实世界病例)上的评估表明,MACD 显著提升了开源大模型的初诊准确率,表现超越了权威知识库,并展现出强大的“人机协同”潜力。
MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM
arXiv ID: 2509.20067
Primary Subject: Artificial Intelligence (cs.AI)
Authors: Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
Timeline: Submitted on September 24, 2025; Last revised on August 24, 2026 (v5)
📋 总结 (Summary)
Large Language Models (LLMs) show immense promise in medical diagnosis, yet current prompting and multi-agent systems typically focus on isolated inferences without accumulating reusable clinical experience. To bridge this gap, this paper introduces MACD (Multi-Agent Clinical Diagnosis), a framework enabling LLMs to self-learn clinical knowledge through a structured multi-agent pipeline—summarizing, refining, and applying insights much like human physicians.
Additionally, the authors propose a MACD-human collaborative workflow combining iterative LLM consultations, a judge agent, and human oversight. Evaluated on the newly constructed MIMIC-MACD cohort (4,390 real-world cases across seven diseases), MACD significantly enhances open-weight LLMs' primary diagnostic accuracy, outperforming authoritative knowledge bases and demonstrating powerful human-AI synergy.
大语言模型(LLM)在辅助医学诊断方面显示出了广阔的前景,基于提示词的方法为其能力提升提供了一种灵活且可部署的途径。然而,现有的提示工程和多智能体方法往往侧重于优化单一推理,较少关注从临床实践中积累可复用的经验,从而限制了它们的现实适用性。
为了解决这一问题,本研究提出了一种新颖的多智能体临床诊断(MACD)框架,该框架允许大模型通过多智能体流水线自主学习临床知识,实现诊断见解的总结、提炼和应用,从而模仿人类医师的专业成长过程。我们进一步将其扩展到MACD-human协同工作流中,其中多个基于LLM的诊断智能体参与迭代会诊,并在未能达成一致的病例中由裁判智能体和人类监督提供支持。
研究所构建的MIMIC-MACD队列包含来自7种疾病的4,390个真实世界患者病例,其中包括1,314个用于知识学习的病例和3,076个用于评估的保留病例。在各种开源大模型中,MACD显著提升了初诊准确率,相比已建立的权威知识平均提升了11.6个百分点,同时缩小了开源模型与顶尖专有大模型之间的性能差距。此外,MACD-human工作流在纯文本病例 vignettes(病例小品)上的表现比纯医师诊断提升了18.3个百分点,证明了人机协同的协同潜力。这项工作因此提出了一个可扩展的自学习范式,弥合了LLM内在知识与现实临床实践需求之间的鸿沟,向着可靠、可解释且可部署的AI辅助诊断迈出了重要一步。
🧠 摘要 (Abstract)
Large language models (LLMs) have shown promise in supporting medical diagnosis, with prompting-based methods offering a flexible and deployable means of capability enhancement. However, existing prompt engineering and multi-agent approaches often focus on optimizing single inferences, paying less attention to the accumulation of reusable experience from clinical practice, constraining their real-world applicability.
To address this, this study proposes a novel Multi-Agent Clinical Diagnosis (MACD) framework, which allows LLMs to self-learn clinical knowledge via a multi-agent pipeline that summarizes, refines, and applies diagnostic insights, mirroring the professional development of human physicians. We further extend it to a MACD-human collaborative workflow, where multiple LLM-based diagnostician agents engage in iterative consultations, supported by a judge agent and human oversight for cases where agreement is not reached.
The MIMIC-MACD cohort comprising 4,390 real-world patient cases across seven diseases is constructed, including 1,314 cases for knowledge learning and 3,076 held-out cases for evaluation. Across diverse open-weight LLMs, MACD significantly improves primary diagnostic accuracy, achieving an average improvement of 11.6 percentage points over established authoritative knowledge, while narrowing the performance gap between open-weight models and state-of-the-art LLMs. Furthermore, the MACD-human workflow yields an 18.3-percentage-point improvement over physician-only diagnosis on text-only vignettes, demonstrating the synergistic potential of human-AI collaboration. This work thus presents a scalable self-learning paradigm that bridges the gap between the intrinsic knowledge of LLMs and the demands of real-world clinical practice, advancing towards a reliable, interpretable, and deployable AI-assisted diagnosis.
大语言模型(LLM)在辅助医学诊断方面显示出了广阔的前景,基于提示词的方法为其能力提升提供了一种灵活且可部署的途径。然而,现有的提示工程和多智能体方法往往侧重于优化单一推理,较少关注从临床实践中积累可复用的经验,从而限制了它们的现实适用性。
为了解决这一问题,本研究提出了一种新颖的多智能体临床诊断(MACD)框架,该框架允许大模型通过多智能体流水线自主学习临床知识,实现诊断见解的总结、提炼和应用,从而模仿人类医师的专业成长过程。我们进一步将其扩展到MACD-human协同工作流中,其中多个基于LLM的诊断智能体参与迭代会诊,并在未能达成一致的病例中由裁判智能体和人类监督提供支持。
研究所构建的MIMIC-MACD队列包含来自7种疾病的4,390个真实世界患者病例,其中包括1,314个用于知识学习的病例和3,076个用于评估的保留病例。在各种开源大模型中,MACD显著提升了初诊准确率,相比已建立的权威知识平均提升了11.6个百分点,同时缩小了开源模型与顶尖专有大模型之间的性能差距。此外,MACD-human工作流在纯文本病例 vignettes(病例小品)上的表现比纯医师诊断提升了18.3个百分点,证明了人机协同的协同潜力。这项工作因此提出了一个可扩展的自学习范式,弥合了LLM内在知识与现实临床实践需求之间的鸿沟,向着可靠、可解释且可部署的AI辅助诊断迈出了重要一步。
🔗 链接与资源 (Links & Resources)
- Paper Access: arXiv:2509.20067 | PDF Link | HTML Version
- Dataset: MIMIC-MACD Cohort (4,390 clinical cases)
- Citations & References: Google Scholar | Semantic Scholar | NASA ADS
(License Reference: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International)
