arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22248cs.CLcs.AIcs.IR

检查点并不足够:CoSLR——一个用于系统文献综述的人机协同系统中的信任校准

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

MD Aidul Islam, Malik Abdul Sami, Muhammad Waseem, Zeeshan Rasheed, Kai-kristian Kemell, Zheying Zhang, Pekka Abrahamsson

首次发表
浏览论文内容

中文总结 AI 辅助

针对系统文献综述中AI生成内容可能未经核实即被信任的问题,提出带强制人工检查点的CoSLR系统;实验显示其可用性良好,但34.9%参与者仍过度信任AI输出,表明需通过界面设计校准信任。

中文摘要 AI 辅助

系统文献综述(SLRs)对于循证研究至关重要,但依然耗时费力,要求研究人员在规划、筛选、分析和报告阶段管理大量文献。大型语言模型(LLMs)现在能够生成流畅、结构良好的综述文本,这使得区分经研究人员验证的综合内容与仅看似权威的综合内容变得困难。这增加了未经核实的AI生成综合内容以系统综述的可信度进入学术记录的风险。我们提出了CoSLR,一个基于大型语言模型和检索增强生成(RAG)的模块化三阶段流水线,支持SLR工作流的人机协同多智能体系统,并在生成输出与其被接受之间设置了明确、强制的人工检查点。在一项基于调查的、有63名参与者参与的研究中,该系统获得了积极评价:63名参与者中有27名(42.9%)对其可用性给予高度评价,表明强制检查点并未以牺牲可用界面为代价。然而,检查点只有在研究人员使用它进行核实时才能保障综述质量:63名参与者中有22名(34.9%)报告称,在与系统进行短暂交互后,他们就会信任AI生成的摘要和报告,而无需额外的人工核查。这些发现表明,人机协同可以支持文献综述工作,但人工监督的有效性取决于用户是否愿意行使监督。这是一个校准问题,界面设计必须直接解决,而不能假设其自动解决。

英文摘要

Systematic Literature Reviews (SLRs) are essential for evidence-based research but remain time-consuming, requiring researchers to manage large volumes of publications across planning, screening, analysis, and reporting. Large language models (LLMs) can now produce fluent, well-structured review text, which makes it difficult to distinguish synthesis that was verified by a researcher from synthesis that merely appears authoritative. This raises the risk that unverified AI-generated synthesis enters the scholarly record carrying the credibility of a systematic review. We present CoSLR, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation (RAG), and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance. In a survey-based study with 63 participants, the system was received positively: 27 of 63 participants (42.9 percent) rated its usability highly, indicating that the mandatory checkpoints did not come at the cost of a workable interface. However, a checkpoint safeguards the review only if researchers use it to verify: 22 of 63 participants (34.9 percent) reported that they would trust AI-generated summaries and reports without additional human checking after only a short interaction with the system. These findings indicate that Human-AI collaboration can support literature review work, but that the effectiveness of human oversight depends on whether users are willing to exercise it. This is a calibration problem that interface design must address directly, not assume.

发表机构

  • Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

↑