arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21374cs.AI

LitReview Arena:基于对战式同行评审平台的文献综述智能体评估

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出对战式文献综述评估平台LitReview Arena,收集约3000条专家判断发现现有最强LLM综述系统仅23.0%胜率,校准后的评估器LitJudge将一致性提升至0.78。

中文摘要 AI 辅助

文献综述对科学进步至关重要,但严格评估自动生成的综述仍存在困难,因为研究效用的许多方面依赖专家判断而非参考重叠指标。我们推出LitReview Arena,这是一款专为文献综述质量定制的对战式评估平台:具备AI论文写作经验的领域专家会对比匿名草稿,被匹配到其专业领域内的主题,并针对五个文献综述专属标准提供维度化结果。通过该协议,我们收集了约3000条专家判断,每条包含五个维度结果,结果显示,当前最强的系统在与人类草稿的决定性对决中,整体效用胜率仅为23.0%;而Sonar Deep Research等智能体式大语言模型(LLM)的表现远超基础语言模型,提升幅度超过60%。我们进一步发现,现有的LLM作为评审者(LLM-as-a-judge)方法与人类专家存在显著不一致(斯皮尔曼相关系数rho=0.467),尤其在论文结构、研究建议等侧重综合的标准上。利用收集的偏好数据,我们构建了经专家校准的评估器LitJudge,其与人类专家的一致性提升至斯皮尔曼相关系数rho=0.78,可与专家间一致性相媲美;代码和数据已公开于此httpsURL。

英文摘要

Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.

补充信息

↑