文章背景与核心概要
验证临床护理是否遵循循证医学指南是机器学习与形式化推理交叉领域的一个复杂挑战。在安全攸关的医疗环境中,单纯依赖纯神经网络或纯符号学范式都存在巨大的潜在风险。为此,研究人员提出了一种专家引导的神经符号学流水线,它将大语言模型(LLMs)严格限制在语义归一化任务中——即把非结构化的药物和微生物描述映射到统一的临床词汇表中。归一化之后,由Sugeno模糊推理系统对这些标准化事件进行处理。该模糊层编码了八项《拯救脓毒症运动》(Surviving Sepsis Campaign)集束化治疗规则,用\([0, 1]\)范围内的分级评分替代了僵硬的二元判断。
通过对MIMIC-IV v3.1数据集中的2,438个脓毒症病例进行评估,该流水线成功揭示了多项关键的临床洞察: * 抗生素使用时机: 被识别为医疗护理中最严重的薄弱环节(平均评分为 \(0.24\),仅有 \(13\%\) 的患者在一小时窗口期内用药)。 * 第1小时集束化治疗表现: 显示出普遍的执行不力(平均评分达 \(36.7\%\))。 * 乳酸监测流失率: 揭示了高乳酸管理过程中存在 \(51\%\) 的流失率。 * ICU住院时长: 突出了不同依从性组别在住院时长上的描述性差异(\(3.8\) 天对比 \(5.1\) 天)。
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights
Summary
验证临床护理是否遵循循证医学指南是机器学习与形式化推理交叉领域的一个复杂挑战。在安全攸关的医疗环境中,单纯依赖纯神经网络或纯符号学范式都存在显著风险。
Verifying whether clinical care adheres to evidence-based protocols represents a complex challenge at the intersection of machine learning and formal reasoning. Relying purely on neural or symbolic paradigms in safety-critical medical environments presents significant risks.
为了解决这一问题,研究人员提出了一种专家引导的神经符号学流水线,它将大语言模型(LLMs)严格限制在语义归一化——将非结构化的药物和微生物描述映射到统一的临床词汇表。在归一化之后,Sugeno 模糊推理系统对这些标准化事件进行处理。该模糊层编码了八条《拯救脓毒症运动》集束化治疗规则,用 \([0, 1]\) 范围内的分级评分取代了僵硬的二元判断。
To address this, researchers propose an expert-guided neuro-symbolic pipeline that strictly constrains Large Language Models (LLMs) to semantic normalization—mapping unstructured drug and microbiology descriptions onto a unified clinical vocabulary. Following normalization, a Sugeno fuzzy inference system processes these standardized events. This fuzzy layer encodes eight Surviving Sepsis Campaign bundle rules, replacing rigid binary judgments with graded scores ranging from \([0, 1]\).
通过对 MIMIC-IV v3.1 数据集中的 2,438 个脓毒症病例进行评估,该流水线成功揭示了关键的临床洞察: * 抗生素使用时机: 被识别为护理中最严重的崩溃点(平均评分为 \(0.24\),仅有 \(13\%\) 在一小时窗口内给药)。 * 第1小时表现: 显示出广泛的执行不力(平均评分为 \(36.7\%\))。 * 乳酸监测流失: 揭示了高乳酸管理中存在 \(51\%\) 的流失率。 * ICU 住院持续时间: 突出了不同依从性组别在住院时长上的描述性差异(\(3.8\) 天对比 \(5.1\) 天)。
Evaluated on 2,438 sepsis episodes from the MIMIC-IV v3.1 dataset, the pipeline successfully uncovered critical clinical insights: * Antibiotic Timing: Identified as the most severe breakdown in care (mean score of \(0.24\), with only \(13\%\) administered within the one-hour window). * Hour-1 Performance: Showed widespread underperformance (mean score of \(36.7\%\)). * Lactate Drop-off: Revealed a \(51\%\) drop-off in elevated-lactate management. * ICU Stay Duration: Highlighted descriptive differences in length of stay across compliance groups (\(3.8\) days versus \(5.1\) days).
Document Metadata
| 元数据字段 | 详情 |
|---|---|
| arXiv ID | arXiv:2608.13617 [cs.AI] |
| 主要学科 | 人工智能 (cs.AI) |
| 次要学科 | 符号计算 (cs.SC) |
| 作者 | Himanshu Tripathi, Kaushik Roy, Subash Neupane, Shahram Rahimi |
| 提交日期 | 2026年8月12日 |
| 会议接收 | 已被 NeSy 2026 接收 (会议链接) |
| DOI | 10.48550/arXiv.2608.13617 |
| 许可证 | 知识共享署名 4.0 |
Metadata Field Details arXiv ID arXiv:2608.13617[cs.AI]Primary Subject Artificial Intelligence ( cs.AI)Secondary Subjects Symbolic Computation ( cs.SC)Authors Himanshu Tripathi, Kaushik Roy, Subash Neupane, Shahram Rahimi Submission Date August 12, 2026 Conference Acceptance Accepted in NeSy 2026 (Conference Link) DOI 10.48550/arXiv.2608.13617License Creative Commons Attribution 4.0
Abstract
验证临床护理是否遵循循证方案是一个天然的神经符号学问题,然而安全攸关的环境使得任何单一范式都难以单独胜任。我们提出了一种专家引导的流水线,该流水线严格将大语言模型限制在语义归一化任务中,将凌乱的药物和微生物学字符串映射到固定的临床词汇表上,同时由 Sugeno 模糊推理系统对归一化的事件进行推理。该模糊层编码了八条《拯救脓毒症运动》集束化治疗规则,并用 [0,1] 范围内的分级评分替代了二元判断。将其应用于 MIMIC-IV v3.1 的 2,438 个脓毒症病例后,发现抗生素使用时机是最关键的崩溃点(平均 0.24,13% 在一小时内)、第1小时执行不力(平均 36.7%)、高乳酸监测流失率达 51%,以及不同依从性组间 ICU 住院时长的描述性差异(3.8 天对比 5.1 天)。
Verifying whether clinical care follows evidence-based protocols is a natural neuro-symbolic problem, yet the safety-critical setting defeats either paradigm alone. We present an expert-guided pipeline that constrains a large language model strictly to semantic normalization, mapping messy drug and microbiology strings onto a fixed clinical vocabulary, while a Sugeno fuzzy inference system reasons over the normalized events. The fuzzy layer encodes eight Surviving Sepsis Campaign bundle rules and replaces binary judgments with graded scores in [0,1]. Applied to 2,438 MIMIC-IV v3.1 sepsis episodes, it surfaces antibiotic timing as the most critical breakdown (mean 0.24, 13% within one hour), Hour-1 underperformance (mean 36.7%), a 51% elevated-lactate drop-off, and descriptive differences in ICU stay across compliance groups (3.8 versus 5.1 days).
Full-Text and Resource Links
- 获取论文: 查看 PDF | HTML(实验性) | TeX 源码
- 外部索引: Google Scholar | Semantic Scholar | NASA ADS
- Access Paper: View PDF | HTML (Experimental) | TeX Source
- External Indices: Google Scholar | Semantic Scholar | NASA ADS