文章背景与核心概要
尽管临床医生经常将思维链(CoT)解释视为大语言模型(LLM)具备真正医疗推理能力的证据,但这些可见的推理链是否真的驱动了模型的决策,目前尚缺乏充分验证。本文作者引入了一套包含 30 个算子的严格医疗扰动审计框架,在四个医疗问答基准上对 14 个 LLM 进行了测试。
研究揭示了一个令人担忧的脱节现象:LLM 往往通过“装饰性”推理得出正确的诊断——这意味着生成的推理链通常无法反映临床编辑的逻辑影响,且完全移除 CoT 提示词也并未导致模型准确率下降。该审计框架为判断医疗 AI 的推理是真正的忠实性表达,还是仅仅流于形式的事后文档,建立了一个亟需的评估标尺。
Right Diagnoses, Decorative Reasoning: A Perturbation Audit of Medical Chain-of-Thought
arXiv: 2608.24790 [cs.AI]
Submitted: August 25, 2026
Authors: Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long
📌 Executive Summary
While clinicians often rely on Chain-of-Thought (CoT) rationales as evidence of true medical reasoning in Large Language Models (LLMs), this study investigates whether those visible reasoning chains actually drive model decisions. The authors introduce a rigorous medical perturbation audit featuring a 30-operator battery to test 14 LLMs across four medical QA benchmarks.
Their findings reveal a troubling disconnect: LLMs frequently arrive at correct diagnoses via "decorative" reasoning—meaning the generated chains often do not reflect the logical impact of clinical edits, and removing CoT prompting entirely fails to degrade accuracy. This audit framework establishes a much-needed yardstick for determining whether medical AI reasoning is genuinely faithful or merely post-hoc documentation.
👥 Authors & Affiliations / 作者与机构
临床医生将思维链(CoT)解释视为医疗推理的证据,但这种可见的推理链是否真正发挥了该作用却极少被测试。通用领域的 CoT 忠实度探测忽略了临床成本,而医疗 LLM 的评估则将推理链视为黑盒。
- Mengzhu Xu
- Jifan Gao
- Xia Jiang
- Yaoxin Wu
- Xi Long
📖 Abstract / 摘要
我们通过医疗扰动审计填补这一空白:包含 30 个算子的算子库利用具有临床动机的操作(严重程度反转、否定翻转、人口统计学对调、证据消融)对推理链和问题进行编辑,并结合“推理链更新与答案翻转”的联合分析,按失效模式对每个模型进行分类。
Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely tested. General-domain CoT-faithfulness probes ignore clinical cost, and medical LLM evaluations treat the chain as a black box.
We close this gap with a medical perturbation audit: a 30-operator battery edits both the chain and the question with clinically motivated operators (severity reversal, negation flip, demographic swap, evidence ablation), paired with a chain-update times answer-flip joint analysis that classifies each model by its failure mode.
应用到四个医疗问答基准上的 14 个 LLM 时,三个独立测试汇聚出以下结论: * 链解耦率(CDR)——即推理链未记录编辑、且答案未翻转的情况——在面向临床有意义的破坏性编辑的面板整体中高达 72.9%。 * 推理链损坏不会对准确率产生影响。 * 移除 CoT 提示词不会降低准确率。
Applied to 14 LLMs on four medical QA benchmarks, three independent tests converge: * The Chain-Decoupling Rate (CDR)—where the chain does not register the edit and the answer does not flip—is 72.9% panel-wide on clinically meaningful destructive edits. * Chain corruption leaves accuracy unchanged. * Removing CoT prompting does not reduce accuracy.
两名经委员会认证的临床医生对 \(N=197\) 个受扰动的问题进行了重新标注,发现 98.5% 的金标准(gold standard)依然具备辩护性。这种模式在医疗微调、推理微调以及不同规模的模型中均普遍存在。在无法获取推理链文本的闭源层级中,答案端的信号与这种解耦现象保持一致。我们的框架和 CDR 提供了一个可复用的标尺,用于审计医疗 CoT 究竟是忠实的推理还是仅仅流于形式的文档。
Two board-certified clinicians re-annotate \(N=197\) perturbed questions, finding that 98.5% leave the gold standard defensible. This pattern holds across medical and reasoning fine-tuning and scale. On the closed-source tier, where the chain text is unavailable, the answer-side signals are consistent with the same decoupling. Our framework and CDR provide a reusable yardstick for auditing whether medical CoT is faithful or merely documentation.
🔗 Additional Resources & Links / 其他资源与链接
- View PDF: Download PDF
- HTML Version: arXiv HTML (experimental)
- TeX Source: Source Files
- License: Creative Commons Attribution 4.0

- External Citations:
- NASA ADS
- Google Scholar
- Semantic Scholar
- View PDF: Download PDF
- HTML Version: arXiv HTML (experimental)
- TeX Source: Source Files
- License: Creative Commons Attribution 4.0
- External Citations:
- NASA ADS
- Google Scholar
- Semantic Scholar