arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在概念复杂的范围综述中评估人工与大语言模型(LLM)筛选工作流:召回率-工作量权衡与运行间一致性

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig

arXiv 2608.26885首次发表:更新:

发表机构

Julius-Maximilians-Universität Würzburg; Institute of Psychology(维尔茨堡大学; 心理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比人工与LLM筛选工作流,发现LLM筛选性能依赖工作流,Gemini 3.1文件批处理召回率最高,GPT-5.4文件批处理运行间存在差异,高召回率任务中LLM更适合人工监督工作流。

AI 中文摘要

背景:大语言模型(LLM)正越来越多地被用于证据合成中的筛选工作,其中假阴性会在全文评估前移除相关研究。我们在一项嵌入概念复杂的范围综述的预注册研究中,比较了人工与LLM的标题-摘要筛选工作流。方法:在仅标题的保守筛选后,1131条记录由1名综述负责人、4名经过训练的助手(筛选非重叠子集),以及7次使用不同模型和处理配置的完整LLM运行(包括名义上相同的重复运行)进行筛选。我们比较了保留的工作量、针对316条已验证合格记录的操作召回率、一致性、运行间一致性及程序负担。由于仅对父综述中推进并评估的记录验证了合格性,召回率估计为操作层面的。结果:没有任何工作流能恢复所有已验证合格记录。人工工作流与2次GPT-5.4文件批处理运行保留了42.2%-45.0%的记录,同时达到82.3%-82.9%的召回率;Gemini 3.1文件批处理达到最高召回率(83.9%),但保留了56.7%的记录。一次性处理配置比对应的文件批处理配置恢复的合格记录更少。2次名义上相同的GPT-5.4文件批处理运行在91.7%的记录上达成一致,但在94条记录上存在差异,其中包括仅被其中一次运行保留的29条已验证合格记录。讨论:LLM筛选性能取决于所实施的工作流,而非仅模型本身。因此,处理配置、工作量、记录层面差异及人工-LLM决策整合是部署系统的重要属性。对于高召回率任务,LLM更适合经过验证、可审计、人工监督的工作流,而非自主排除。

英文摘要

Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑