SciLitBench:面向大语言模型驱动的系统文献综述的基准与设计原则
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
- Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SciLitBench是一个多阶段系统综述基准,涵盖筛选与数据提取,发现明确标准提升筛选性能,但提取可靠性存在边界,为LLM辅助证据综合提供可复现评估资源。
AI中文摘要:
系统综述需要在数千条记录上进行持续的人类判断,然而现有对大语言模型(LLM)的评估通常孤立地考察综述的各个阶段。我们引入了SciLitBench,一个多阶段基准,涵盖标题与摘要筛选、全文筛选以及模式引导的数据提取,包含42,981条检索记录、1,012篇全文以及888篇纳入论文的标注。在来自六个模型家族的22个开放权重大语言模型中,明确的纳入和排除标准使标题与摘要筛选的$F_2$分数提高了28.8%,而研究者撰写的理由说明使全文筛选提高了15%。数据提取揭示了不同的可靠性机制:性能从出版年份的0.97准确率下降到计算方法上的0.37 Jaccard重叠度,而最强的模型仅恢复了30%的标注评估证据和25%的局限性。SciLitBench确定了高召回率筛选与证据完整提取之间的实际边界,并为评估LLM辅助的证据综合提供了可复现的资源。
英文摘要:
Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30\% of annotated evaluation evidence and 25\% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.