LitReview Arena:基于对战式同行评审平台的文献综述智能体评估
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
浏览论文内容
中文总结 AI 辅助
该研究推出对战式文献综述评估平台LitReview Arena,收集约3000条专家判断发现现有最强LLM综述系统仅23.0%胜率,校准后的评估器LitJudge将一致性提升至0.78。
中文摘要 AI 辅助
文献综述对科学进步至关重要,但严格评估自动生成的综述仍存在困难,因为研究效用的许多方面依赖专家判断而非参考重叠指标。我们推出LitReview Arena,这是一款专为文献综述质量定制的对战式评估平台:具备AI论文写作经验的领域专家会对比匿名草稿,被匹配到其专业领域内的主题,并针对五个文献综述专属标准提供维度化结果。通过该协议,我们收集了约3000条专家判断,每条包含五个维度结果,结果显示,当前最强的系统在与人类草稿的决定性对决中,整体效用胜率仅为23.0%;而Sonar Deep Research等智能体式大语言模型(LLM)的表现远超基础语言模型,提升幅度超过60%。我们进一步发现,现有的LLM作为评审者(LLM-as-a-judge)方法与人类专家存在显著不一致(斯皮尔曼相关系数rho=0.467),尤其在论文结构、研究建议等侧重综合的标准上。利用收集的偏好数据,我们构建了经专家校准的评估器LitJudge,其与人类专家的一致性提升至斯皮尔曼相关系数rho=0.78,可与专家间一致性相媲美;代码和数据已公开于此httpsURL。
英文摘要
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.