跳转至

利用少样本学习与大语言模型分析科学文献中不同生物学性别的血压变化

文章背景与核心概要

本研究探讨了如何利用大语言模型(LLM)和自然语言处理(NLP)技术,从科学文献中自动化提取血压(BP)数据,并特别关注生物学性别等人口统计学差异。传统的血压标准往往忽略了个体的人口统计学因素,从而降低了诊断的可靠性。为了解决这一问题,研究人员开发了一个基于Solr的PubMed搜索引擎,筛选并人工审核了213篇文献的子集(其中90篇报告了基于性别的血压值),并测试了少样本和零样本学习方法(LLaMA3和GPT-3.5),以提取血压均值、标准差及生物学性别。

随后,提取出的数据被用于生成热力图和等高线图,从而深入分析不同生物学性别之间的血压分布变化。这项工作展示了人工智能在推进个性化医疗和临床数据自动化提取方面的巨大潜力。


license icon

Summary

本研究探讨了如何利用大语言模型(LLM)和自然语言处理(NLP)技术,从科学文献中自动化提取血压(BP)数据,并特别关注生物学性别等人口统计学差异。传统的血压标准往往忽略了个体的人口统计学因素,从而降低了诊断的可靠性。为了解决这一问题,研究人员开发了一个基于Solr的PubMed搜索引擎,筛选并人工审核了213篇文献的子集(其中90篇报告了基于性别的血压值),并测试了少样本和零样本学习方法(LLaMA3和GPT-3.5),以提取血压均值、标准差及生物学性别。随后,提取出的数据被用于生成热力图和等高线图,以分析不同生物学性别之间的血压分布变化。

This study explores the automation of extracting blood pressure (BP) data from scientific literature using Large Language Models (LLMs) and Natural Language Processing (NLP) techniques, with a specific focus on demographic distinctions such as biological sex. Traditional BP standards often overlook individual demographic factors, reducing diagnostic reliability. To address this, the researchers developed a Solr-based search engine for PubMed, curated a manually reviewed subset of 213 articles (including 90 reporting sex-based BP values), and tested few-shot and zero-shot learning methods (LLaMA3 and GPT-3.5) to extract BP means, standard deviations, and biological sex. The extracted data was then used to generate heatmaps and contour plots to analyze BP distribution variations across biological sexes.


Metadata

  • arXiv ID: arXiv:2402.01826 [cs.CL]
  • Subjects: 计算与语言 (cs.CL); 人工智能 (cs.AI)
  • Authors: Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker
  • Submitted On: 2024年2月2日 (v1); 最后修订于 2026年8月14日 (v2)
  • Status: 已被期刊 Computers in Biology and Medicine 接受
  • Related DOI: 10.1016/j.compbiomed.2025.111128

Abstract

当前的血压(BP)技术和标准建立于数十年前,这些标准至今仍在世界范围内沿用,且通常在没有根据性别和年龄等个体人口统计学因素调整血压读数的情况下使用。虽然这些标准提供了有用的指南并有助于识别高危患者,但由于缺乏对人口统计学因素的考量,它们在诊断上并非完全可靠。本研究旨在评估利用大语言模型(LLM)从科学文献中自动提取血压相关信息的 Så 可行性,重点关注血压分布中基于生物学性别的差异。

Current blood pressure (BP) technologies and standards were established decades ago, and these standards are still used worldwide today, often without adjusting BP readings for individual demographic factors such as sex and age. While these standards provide useful guidelines and help identify at-risk patients, they are not fully reliable for diagnosis due to the lack of demographic considerations. This study aims to assess the feasibility of using large language models (LLMs) for the automated extraction of BP-related information from the scientific literature, with a focus on biological sex-based distinctions in BP distributions.

我们采用自然语言处理(NLP)方法,从文献中提取出区分生物学性别的血压值的均值和标准差。我们开发了一个基于Solr的搜索引擎,从PubMed中检索包含血压相关关键词和生物学性别指标的科学文献。从检索到的文章中,我们创建了一个包含213篇文章的人工审核子集,其中包括90个报告基于生物学性别的血压值案例。我们试验了一种少样本学习方法和两种基于LLM的零样本方法(LLaMA3和GPT-3.5),以提取血压值的均值、标准差以及相关的生物学性别。基于自动提取的信息,我们生成了热力图和等高线图,以研究不同生物学性别之间血压值的变化。

We employed natural language processing (NLP) methods to extract the means and standard deviations of BP values from the literature, distinguishing by biological sex. We developed a Solr-based search engine to retrieve scientific articles containing BP-related keywords and biological sex indicators from PubMed. From the retrieved articles, we created a manually reviewed subset comprising 213 articles including 90 cases that reported BP values based on biological sex. We experimented with one few-shot learning method and two zero-shot LLM-based methods—LLaMA3 and GPT-3.5—to extract the mean and standard deviations of BP values, and the associated biological sex. Based on the automatically-extracted information, we generated heatmaps and contour plots to study the variations of BP values across biological sex.