arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用人类与大语言模型(LLM)的分歧改进基于核查表的质量评估

Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

Timo van der Kuil, Bruno Messina Coimbra, Mirjam van Zuiden, Robert A. Bagheri, Rens van de Schoot, Klaas Dieleman, Berend Greijn, Stefan Houkes, Sebastiaan Rodenhuis, Elizabeth M. Grandfield

arXiv 2608.20385首次发表:更新:

发表机构

Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究LLM能否近似基于核查表评估的人类判断,发现人类与LLM的分歧多源于核查表的歧义性和条件性标准,修订此类条目可提升评估一致性,分析分歧还能助力研究综合工作流的迭代改进。

AI 中文摘要

系统综述依赖于对纳入研究的质量评估,该过程耗时且易受核查表标准歧义的影响。尽管大语言模型(LLM)为支持这些任务提供了机会,但评估核查表通常被视为固定输入,其设计如何影响与专家判断的一致性仍不明确。因此,本研究调查两个问题:(1)LLM是否能近似基于核查表评估的人类判断;(2)人类与LLM的分歧模式是否可用于识别和改进歧义核查表条目。本研究采用《潜轨迹研究报告指南》(GRoLTS)核查表,在三个研究主题和两个核查表版本中,比较LLM生成的评估结果与专家标注结果,通过条目级准确率、 chance-corrected 一致性( chance-corrected agreement)及研究级排序的保留情况评估一致性。研究发现,不同核查表条目的性能差异显著,歧义性和条件性标准会产生最大分歧;修订这些条目可提升原始一致性和 chance-corrected 一致性。尽管条目级误分类仍存在,但当保留高一致性条目时,LLM生成的评分通常能保留研究的相对排序。这些结果表明,可靠的LLM辅助评估不仅取决于模型选择,还取决于核查表设计;研究结论指出,分析人类与LLM的分歧可帮助识别有问题的核查表条目,支持研究综合工作流的迭代改进。

英文摘要

Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treated as fixed inputs, and it remains unclear how their design affects agreement with expert judgments. Therefore, we investigate (1) whether LLMs can approximate human judgments in checklist-based appraisal and (2) whether patterns of human-LLM disagreement can be used to identify and improve ambiguous checklist items. Using the Guidelines for Reporting on Latent Trajectory Studies (GRoLTS) checklist, we compare LLM-generated assessments with expert annotations across three research topics and two checklist versions. Agreement is assessed using item-level accuracy, chance-corrected agreement, and preservation of study-level rank ordering. We find that performance varies substantially across checklist items, with ambiguous and conditional criteria producing the greatest disagreement. Revising these items improves both raw and chance-corrected agreement. Although item-level misclassifications persist, LLM-generated scores often preserve the relative ranking of studies when high-agreement items are retained. These results indicate that reliable LLM-assisted appraisal depends not only on model choice but also on checklist design. The findings suggest that analyzing human-LLM disagreement can help identify problematic checklist items and support the iterative improvement of research synthesis workflows.

Comments31 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑