跳转至

文章背景与核心概要

在X射线透视引导下对血管内介入手术中使用的细长医疗器械(如导管和导丝)进行准确的视觉分析,对于保障手术安全至关重要。然而,基于学习的自动化方法面临着重大障碍,包括结构复杂性、数据稀缺性以及严格的患者隐私法规(这些法规限制了集中式的机构间模型训练)。

本篇博士论文引入了一个全面的结构感知联邦学习框架,旨在实现无需集中患者数据或进行大量人工标注的协同导管与导丝分析。该研究在真实动物数据集和体模数据集上进行了验证,主要取得了四项方法论贡献:构建了大规模基准数据集 CathAction;提出了能将掩码转化为符号距离图的形状敏感损失函数;设计了结合形状敏感损失与投影梯度下降的联邦学习方法;以及开发了结合结构监督与领域自适应重建目标的结构感知扩散框架,有效解决了数据稀缺条件下的分割难题。


Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting

Author: Chayun Kongtongvattana (PhD Thesis, University of Liverpool)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
arXiv: arXiv:2609.06876 [cs.CV]
Submitted: September 6, 2026 (163 pages)

Author: Chayun Kongtongvattana (PhD Thesis, University of Liverpool)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
arXiv: arXiv:2609.06876 [cs.CV]
Submitted: September 6, 2026 (163 pages)


📌 Summary

Accurate visual analysis of thin medical instruments—such as catheters and guidewires—under X-ray fluoroscopy is critical for the safety of endovascular procedures. However, learning-based automation faces major obstacles, including structural complexity, data scarcity, and strict patient privacy regulations that prevent centralized institutional training.

This PhD thesis introduces a comprehensive structure-aware federated learning framework designed to enable collaborative catheter and guidewire analysis without centralizing patient data or requiring massive manual annotations. The research is validated across real-animal and phantom datasets, yielding four primary methodological contributions:

  1. The CathAction Benchmark Dataset: A large-scale dataset featuring over 600,000 annotated frames and 40,000 segmentation masks tailored for catheterisation analysis.
  2. Shape-Sensitive Loss: A novel loss function that transforms masks into signed distance maps evaluated in a structural feature space, improving the Dice coefficient by up to 2.9 points across five backbone architectures.
  3. Federated Learning with Shape-Sensitive Loss & Projected Gradient Descent:
  4. Preserves geometric consistency under heterogeneous client data, outperforming standard federated averaging by up to three points in mean Intersection-over-Union (mIoU) as clients scale from four to eight.
  5. Incorporates adversarial optimization via projected gradient descent, boosting mIoU by over 10 points on real-animal data.
  6. Structure-Aware Diffusion Framework: A generative approach that synthesizes catheter and guidewire video sequences by combining structural supervision with a domain-adaptive reconstruction objective. Integrating synthetic data into federated training raises the Dice score from 44% to 51% under data scarcity conditions across four held-out sites.

Accurative visual analysis of thin medical instruments—such as catheters and guidewires—under X-ray fluoroscopy is critical for the safety of endovascular procedures. However, learning-based automation faces major obstacles, including structural complexity, data scarcity, and strict patient privacy regulations that prevent centralized institutional training.

This PhD thesis introduces a comprehensive structure-aware federated learning framework designed to enable collaborative catheter and guidewire analysis without centralizing patient data or requiring massive manual annotations. The research is validated across real-animal and phantom datasets, yielding four primary methodological contributions:

  1. The CathAction Benchmark Dataset: A large-scale dataset featuring over 600,000 annotated frames and 40,000 segmentation masks tailored for catheterisation analysis.
  2. Shape-Sensitive Loss: A novel loss function that transforms masks into signed distance maps evaluated in a structural feature space, improving the Dice coefficient by up to 2.9 points across five backbone architectures.
  3. Federated Learning with Shape-Sensitive Loss & Projected Gradient Descent:
  4. Preserves geometric consistency under heterogeneous client data, outperforming standard federated averaging by up to three points in mean Intersection-over-Union (mIoU) as clients scale from four to eight.
  5. Incorporates adversarial optimization via projected gradient descent, boosting mIoU by over 10 points on real-animal data.
  6. Structure-Aware Diffusion Framework: A generative approach that synthesizes catheter and guidewire video sequences by combining structural supervision with a domain-adaptive reconstruction objective. Integrating synthetic data into federated training raises the Dice score from 44% to 51% under data scarcity conditions across four held-out sites.