概念复杂的范围综述中人工与大语言模型筛选工作流评估:召回率-工作量权衡与运行一致性
文章背景与核心概要
在科学研究与证据综合(Evidence Synthesis)领域,文献的标题和摘要筛选是至关重要的一步。然而,在概念复杂的范围综述中,由假阴性(漏检)导致的错误排除可能会在全文评估之前就遗漏关键研究。本文由 Nikol Figalová 等研究者撰写,针对这一难题,深入探讨了大语言模型(LLM)与人类审核员在文献筛选阶段的效能表现。
研究团队通过一项预注册研究,对比了人类工作流与多种不同模型及配置的 LLM 工作流(包括文件批处理与一次性处理模式)。研究核心评估了各方案的保留工作量、操作召回率、运行间的一致性以及程序负担。结果表明,没有任何一种工作流能够召回所有经验证的合格记录,LLM 的表现高度依赖于处理配置与人机监督集成,而非单纯依赖模型本身的智能。该研究为高召回率任务中的人机协作筛选工作流提供了重要的实践指导。
评估概念复杂的范围综述中人工与大语言模型筛选工作流:召回率-工作量权衡与运行一致性
Evaluating Human and LLM Screening Workflows in a Conceptually Complex Scoping Review: Recall–Workload Trade-offs and Run-to-Run Consistency
作者: Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig
分类: 计算机科学 > 人工智能 (cs.AI)、人机交互 (cs.HC)、软件工程 (cs.SE)
arXiv ID: arXiv:2608.26885 [cs.AI]
提交时间: 2026年8月27日
链接: 查看 PDF | HTML 版本 | 许可证 
摘要 (Summary)
本研究调查了大语言模型(LLM)在概念复杂的范围综述的标题与摘要筛选阶段,相较于人类审核员的有效性。由于证据综合中的假阴性可能会在全文评估之前无意中排除相关研究,研究人员评估了各种人类和 LLM 工作流。他们的发现揭示了关于操作召回率、保留工作量和运行一致性的关键见解,证明了 LLM 的性能在很大程度上依赖于处理配置和人工监督集成,而不是自主排除。
Summary
This study investigates the effectiveness of Large Language Models (LLMs) compared to human reviewers during the title-and-abstract screening phase of a conceptually complex scoping review. Because false negatives in evidence synthesis can inadvertently exclude relevant studies before full-text evaluation, the researchers evaluated various human and LLM workflows. Their findings reveal critical insights regarding operational recall, retained workload, and run-to-run consistency, demonstrating that LLM performance relies heavily on processing configurations and human-supervised integration rather than autonomous exclusion.
摘要详情 (Abstract)
背景
大语言模型(LLM)正日益被用于证据综合中的筛选工作,在此类任务中,假阴性可能会在进行全文评估前剔除相关研究。我们在一个概念复杂的范围综述中嵌入了一项预注册研究,对比了人类与 LLM 的标题和摘要筛选工作流。
Background
Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review.
方法
在经过保守的仅标题筛选后,1,131 条记录由 1 名综述负责人、4 名分别筛选非重叠子集的受训助理,以及使用不同模型和处理配置(包括一次名义上相同的重复运行)的 7 次完整 LLM 运行进行筛选。我们对比了保留的工作量、针对 316 条经验证合格记录的操作召回率、一致性、运行间一致性以及程序负担。由于合格性仅针对母综述中推进和评估的记录进行验证,因此召回率估计值为操作性估计。
Methods
After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational.
结果
没有任何工作流能够恢复所有经验证的合格记录。人类工作流和两次 GPT-5.4 文件批处理运行保留了 42.2% 至 45.0% 的记录,同时实现了 82.3% 至 82.9% 的召回率。Gemini 3.1 文件批处理实现了最高的召回率(83.9%),但保留了 56.7% 的记录。所有“一次性全部处理”(all-at-once)配置恢复的合格记录均少于相应的“文件批处理”配置。两次名义上相同的 GPT-5.4 文件批处理运行在 91.7% 的记录上达成了一致,但在 94 条记录上存在分歧,其中包括仅由一次运行保留的 29 条经验证合格记录。
Results
No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2–45.0% of records while achieving 82.3–82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run.
讨论
LLM 筛选性能取决于所实施的工作流,而不仅仅是模型本身的身份。因此,处理配置、工作量、记录级变异性以及人机决策集成是部署系统的实质性属性。对于高召回率任务,LLM 更适合采用经过验证、可审计且有人工监督的工作流,而非自主排除。
Discussion
LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
提交历史
- [v1] 2026年8月27日 星期四 09:42:39 UTC (70 KB)
Submission History
- [v1] Thu, 27 Aug 2026 09:42:39 UTC (70 KB)