跳转至

文章背景与核心概要

新生儿呼吸系统疾病仍然是导致新生儿发病和死亡的主要原因之一,在临床诊疗中面临着巨大的诊断挑战。现有的多模态大语言模型(MLLM)由于领域鸿沟(主要基于成人医学影像训练)以及上下文不足(无法将多维临床背景与胸部X光片有效结合),在该领域往往表现不佳。

为了突破这些瓶颈,研究人员推出了 NeoRed——首个专门针对新生儿呼吸系统疾病诊断和报告生成量身定制的多模态大语言模型,并同时开源了两大真实临床数据集:NeoCXR 和 NeoCXR-EV。NeoRed 引入了创新的知识-逻辑-对齐(KLA)框架,通过知识先验注入、诊断逻辑约束和视觉语义对齐三个核心维度规范模型行为,成功弥合了异构临床背景与影像数据之间的差距,在 NeoCXR 数据集上取得了显著优于基线模型的性能表现。


NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

Authors: Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo
ArXiv ID: [arXiv:2609.03527 [cs.AI]]
Submitted: September 3, 2026


Executive Summary

Neonatal respiratory diseases remain a leading cause of morbidity and mortality in newborns, presenting significant diagnostic challenges in clinical settings. Existing Multimodal Large Language Models (MLLMs) fall short in this domain due to two primary bottlenecks: 1. The Domain Gap: Most models are trained predominantly on adult medical imaging and data. 2. Context Insufficiency: Current systems fail to adequately integrate multidimensional clinical contexts alongside chest X-rays (CXRs).

To overcome these barriers, the researchers introduce NeoRed, the first MLLM specifically tailored for neonatal respiratory disease diagnosis and report generation. Alongside the model, they introduce two real-world clinical datasets: NeoCXR and NeoCXR-EV.

To bridge heterogeneous clinical contexts with imaging data, NeoRed utilizes a novel Knowledge-Logic-Alignment (KLA) framework, constraining model behavior via three core perspectives: * Knowledge Prior Injection (KPI): Infuses neonatologist-approved diagnostic priors into multimodal representations to guide disease-specific attention. * Diagnostic Logic Constraint (DLC): Aligns the semantic output of generated reports with formal multimodal diagnostic logic. * Visual Semantic Alignment (VSA): Establishes explicit semantic correspondence between visual features and clinical imaging conclusions.

Key Results

  • Achieves a ROUGE-L score of 53.29% and a Clinical Efficacy F1 score of 65.19% on the NeoCXR dataset, outperforming baseline MLLMs.
  • Maintains competitive performance on standard adult benchmarks (MIMIC-CXR and IU-Xray).

Paper Metadata

Category Details
Primary Subject Artificial Intelligence (cs.AI)
Document Specs 9 pages, 10 figures
Full-Text & Resources View PDF | HTML Version | TeX Source
DOI / Citations 10.48550/arXiv.2609.03527 | Google Scholar | Semantic Scholar