跳转至

文章背景与核心概要

利用大语言模型(LLMs)自动化基于临床指南的决策制定,常常面临可靠性不足、幻觉现象以及缺乏可解释性等挑战。本文探讨了大语言模型及其各类推理策略在自动根据自由文本盆腔超声报告进行卵巢-附件报告和数据系统(O-RADS)分类时的实际效能。

研究团队在包含 310 名女性及 390 个卵巢肿块的回顾性数据集上,评估了 8 种大语言模型以及三种推理策略(隐性知识端到端、规则引导端到端、以及将特征提取与规则分类相分离的基于特征的混合架构)。实验表明,由 Gemini 3.6 Flash 驱动的基于特征的混合架构取得了 99.2%(387/390)的出色准确率,并与专家共识达到了近乎完美的契合度(加权 Kappa 值为 1.00,95% 置信区间:0.99–1.00)。该混合模型不仅显着超越了原始临床报告(准确率 87.7%)和标准端到端大模型策略(准确率在 65.6% 到 95.9% 之间),还有效缓解了原始报告中常见的过度分期(overstaging)倾向,减少了错误分类,为标准化临床决策建立了一种可靠且具可解释性的机制。


Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

arXiv: arXiv:2608.23061 [cs.AI]
DOI: 10.48550/arXiv.2608.23061
Submitted: August 24, 2026
Primary Subject: Artificial Intelligence (cs.AI)

arXiv: arXiv:2608.23061 [cs.AI]
DOI: 10.48550/arXiv.2608.23061
Submitted: August 24, 2026
Primary Subject: Artificial Intelligence (cs.AI)


Abstract Summary

Abstract Summary

Automating clinical guideline-based decision-making using Large Language Models (LLMs) often faces challenges related to reliability, hallucinations, and a lack of interpretability. This study investigated how effectively LLMs and various reasoning strategies can automate Ovarian-Adnexal Reporting and Data System (O-RADS) classification derived from free-text pelvic ultrasound reports.

Automating clinical guideline-based decision-making using Large Language Models (LLMs) often faces challenges related to reliability, hallucinations, and a lack of interpretability. This study investigated how effectively LLMs and various reasoning strategies can automate Ovarian-Adnexal Reporting and Data System (O-RADS) classification derived from free-text pelvic ultrasound reports.

Key Highlights:

  • Methodology: Evaluated 8 LLMs using three reasoning strategies (implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture separating feature extraction from rule-based classification) on a retrospective dataset of 310 women and 390 ovarian masses.
  • Best Performing Setup: The feature-based hybrid architecture powered by Gemini 3.6 Flash achieved an outstanding accuracy of 99.2% (387/390) and near-perfect agreement with expert consensus (weighted kappa = 1.00; 95% CI: 0.99–1.00).
  • Comparison: Outperformed original clinical reports (accuracy 87.7%) and standard end-to-end LLM strategies (accuracy range 65.6% to 95.9%).
  • Impact: The hybrid model successfully mitigated overstaging tendencies found in original reports and reduced misclassification errors, establishing a dependable, interpretable mechanism for standardized clinical decision-making.

Key Highlights:

  • Methodology: Evaluated 8 LLMs using three reasoning strategies (implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture separating feature extraction from rule-based classification) on a retrospective dataset of 310 women and 390 ovarian masses.
  • Best Performing Setup: The feature-based hybrid architecture powered by Gemini 3.6 Flash achieved an outstanding accuracy of 99.2% (387/390) and near-perfect agreement with expert consensus (weighted kappa = 1.00; 95% CI: 0.99–1.00).
  • Comparison: Outperformed original clinical reports (accuracy 87.7%) and standard end-to-end LLM strategies (accuracy range 65.6% to 95.9%).
  • Impact: The hybrid model successfully mitigated overstaging tendencies found in original reports and reduced misclassification errors, establishing a dependable, interpretable mechanism for standardized clinical decision-making.

Authors

Authors

Xiaotong Tan, Chunli Qiu, Xin Liu, Qing Huang, Guangli Zhou, Bo Gao, Xiaoyan Song, Shuyan Wang, Xiuqin Wang, Wufeng Xue, Ruobing Huang, Dong Ni, Guowei Tao, Jun Cheng

Xiaotong Tan, Chunli Qiu, Xin Liu, Qing Huang, Guangli Zhou, Bo Gao, Xiaoyan Song, Shuyan Wang, Xiuqin Wang, Wufeng Xue, Ruobing Huang, Dong Ni, Guowei Tao, Jun Cheng


Additional Information

Additional Information

  • Main Manuscript: 20 pages, 5 figures, 2 tables
  • Supplemental Material: 11 pages, 1 figure, 3 tables
  • Full-Text PDF: View PDF
  • Main Manuscript: 20 pages, 5 figures, 2 tables
  • Supplemental Material: 11 pages, 1 figure, 3 tables
  • Full-Text PDF: View PDF