arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

风格取胜,实质落败:对作为创意生成评判者的大语言模型(LLM)的诊断

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui, Yu Xie, Min Peng, Qianqian Xie, Xiuying Chen, Zhuohan Xie

arXiv 2608.01666首次发表:更新:

AI 中文总结

本研究针对大语言模型作为创意评判者的风格偏差问题,提出SciStyleBench基准及配套的SciStyleExtractor模块,实验验证该模块可有效降低风格偏差、提升实质区分度,为科学创意评估提供了系统性解决方案。

AI 中文摘要

然而,这些评判者究竟是真正评估创意的科学实质,还是受到表面风格呈现的影响,仍是一个悬而未决的问题。为解决该问题,我们提出SciStyleBench,这是一个用于诊断和缓解基于大语言模型的创意评估中风格偏差的统一三组件基准:其一,SciStyleStage是一个三阶段评估环境,在三种设置(无上下文、固定领域上下文、开放领域检索上下文)下对固定科学内容施加可控的风格扰动,涵盖600个科学创意和15种风格变体,每种设置包含9000个评估实例;其二,SciStyleMetrics是一组量化指标,包括风格偏差指数(SBI)、实质识别率(SRR)和对抗胜率(AWR),用于表征风格变化如何影响评分稳定性、实质区分度和排名稳健性;其三,SciStyleExtractor是一个即插即用的评估模块,通过预测风格类型和偏差将呈现风格与科学实质分离,再进行风格条件评估,使我们能够评估风格感知是否可减少风格偏差。在SciStyleBench上的实验表明,直接的大语言模型评判者仍对写作风格敏感,难以区分科学实质。相比之下,SciStyleExtractor将SBI从0.566降至0.501,同时将SRR和AWR从0.504和0.554分别提升至0.759和0.899。这些结果表明,稳健的创意评估需要对风格变化保持不变性,同时不牺牲对科学实质的敏感性。总体而言,SciStyleBench为识别、量化和缓解科学创意评估中的风格偏差提供了系统框架。

英文摘要

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.

CommentsFirst three authors are co-first authors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑