arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30298cs.CL

系统评价中筛选自动化的基准框架

A Benchmark Framework for Screening Automation in Systematic Reviews

Gauransh Kumar, Luciano Marchezan, Guillaume Genois, Kévin Delcourt, Eugene Syriani

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SRBench基准数据集与评估框架,并开发PromptSR工具,以解决系统评价筛选中类别不平衡问题,提升LLM筛选性能评估的可靠性。

中文摘要 AI 辅助

系统评价(SR)对于循证研究至关重要,但其筛选阶段非常耗时且劳动密集。大型语言模型(LLMs)通过协助文章相关性分类,为减少这一工作量提供了有前景的机会。然而,现有的评估方法通常依赖传统指标,这些指标对于高度不平衡的SR筛选可能具有误导性。本文提出了一个包含45,064条标记条目的基准数据集,用于评估LLM在32个精选二级研究中的SR筛选性能。它提出了一个考虑类别不平衡(即SR中排除文章相对于纳入文章的自然普遍性)的评估框架。它还介绍了PromptSR,一个旨在支持基于LLM的筛选的提示实验、实验管理和结果分析的工具。我们还展示了一个用例,演示了SRBench和PromptSR的应用。

英文摘要

Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets. This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.

发表机构

  • Université de Montréal(蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑