arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11559cs.SEcs.AIcs.CL

PRISMA-LLM:人工智能辅助系统评价的经验报告框架

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

Miguel Zabaleta, Baihan Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI辅助系统评价中报告不一致的问题,提出PRISMA-LLM框架,区分实现披露与后果敏感的评估及局限性报告。

中文摘要 AI 辅助

大型语言模型(LLM)和人工智能赋能的软件越来越多地参与系统评价决策,然而审计这些工作流程所需的信息报告并不一致。我们分析了SciLitBench——一个包含888篇综述自动化论文、共14,726条注释的语料库——以刻画方法、综述阶段使用、评估和报告局限性方面的变化。自动化已转向面向LLM和软件的工作流程,包括可能改变证据基础的阶段。自2023年以来,38.0%的软件/产品论文未报告评估,而LLM论文中这一比例为9.3%。报告覆盖率随LLM工作流程复杂性增加而提高,但52%的仅报告正面结果的LLM评估仍报告了未满足的可靠性或性能要求。基于这些模式,我们提出了PRISMA-LLM,一个经验驱动的框架,将实现披露与对后果敏感的评估和局限性报告区分开来。

英文摘要

Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.

发表机构

  • Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑