发表机构
Renmin University of China; Tsinghua University(中国人民大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM评估综述未与人类审稿人对齐的问题,提出SurveyReview基准及SurveyAlign基线,大幅提升自动评估与人类审稿人的对齐程度,为该领域提供参考。
AI 中文摘要
大型语言模型的快速发展已将综述撰写从耗时数月的人工工作转变为自动化流程。随着生成规模扩大,可靠评估成为瓶颈,且大型语言模型(LLM)正越来越多地被用作综述评估者。然而,现有方法大多依赖现成的LLM作为评判者的方法,未与人类审稿人进行系统性对齐,且仍缺乏量化与人类审稿人对齐程度的系统性框架。为解决这一缺口,我们提出SurveyReview,这是一个面向综述评估的、与审稿人对齐的多维度基准及数据集。我们收集并标注了675篇综述论文的1630份审稿报告,通过将自由格式评论转化为四个维度(可读性、批判性、全面性、结构)的分数并辅以支撑理由,构建了真实的同行评审报告。我们进一步发布了标准化的训练/测试拆分及评估协议,用于衡量自动评估者与人类审稿人之间的对齐程度。为验证该基准,我们开发了SurveyAlign,这是一个强基线评估器,通过在我们的标注数据上使用LoRA微调Qwen3-32B,并为知识密集型维度补充外部知识。在测试集上,SurveyAlign在审稿人对齐方面较基于GPT-5.2的提示式评判有大幅提升,在全部四个维度上将平均均方误差(MSE)从2.28降至1.38,平均绝对误差(MAE)从1.15降至0.69。我们的贡献有两点:(1)建立了首个具有可复现评估框架的多维度、与审稿人对齐的综述审稿数据集;(2)开发了一个强基线评估器,大幅提升了与人类审稿人的对齐程度,为未来研究提供了有竞争力的参考。我们的代码和数据可在此httpsURL获取。
英文摘要
The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at https://surveyreview.github.io