arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01905cs.DL

指导大语言模型同行评审者:评分锚点对评审证据与准确性的影响

Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy

  • School of Information, Journalism and Communication(信息、新闻与传播学院)
  • The University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

Judita Preiss, Yunhan Yang

AI总结:

本研究以98份联合健康专业研究成果为对象,对比无指导与神谕指导的LLM评审,发现评分锚点可引导评审理由,且对优势/升级要点的覆盖更可靠,同时提升了评分准确性。

AI中文摘要:

大语言模型(LLM)越来越多地被用于研究质量评估,已有研究探索了其评分准确性和评审理由的合理性。然而,外部评分指导是否会改变生成评审中呈现的证据以及最终评分,目前了解较少。本研究使用98份提交用于REF风格内部评估的联合健康专业研究成果,这些成果配有专业人类评审报告和经裁定的1-4分参考评分。将无指导基线评审与神谕指导评审进行对比,神谕指导评审中提供的评分被设置为取整后的人类参考评分;提取的评估要点用于对比人类与LLM的证据使用情况。通过该设计,神谕指导提升了评分准确性,评分遵循检查显示模型不会简单复制提供的评分。纠正的评分不匹配与生成评审框架的变化相关,表明评分信号能够引导评审理由。该效应具有方向依赖性:LLM评审覆盖人类优势或升级要点的可靠性高于人类劣势或降级要点,其中与专家降级证据的对齐最弱。结果表明,评分指导的评审生成可在评审证据层面以及最终评分层面进行评估。

英文摘要:

Large language models (LLMs) are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales. However, less is known about whether external score guidance changes the evidence presented in the generated review as well as the final score. This study uses 98 Allied Health Professions research outputs submitted for internal REF-style assessment, with specialist human review reports and adjudicated 1-4 reference scores. No-guidance baseline reviews are compared with oracle-guided reviews, where the supplied score is set to the rounded human reference score; extracted evaluation points are used to compare human and LLM evidence use. Using this design, oracle guidance improves scoring accuracy, with score-following checks showing that models do not simply copy the supplied score. Corrected score mismatches are associated with changes in the generated review frame, showing that the score signal can steer review rationales. This effect is direction-dependent: LLM reviews cover human strength or upgrade points more reliably than human weakness or downgrade points, with the weakest alignment for expert downgrade evidence. The results show that score-guided review generation can be evaluated at the level of review evidence, as well as the final score.

↑