arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我应该对我的LLM相关性评判员礼貌吗?语气作为严重性操作点偏移

Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift

Tian Zhang, Meng Li

arXiv 2609.09703首次发表:更新:

AI 中文总结

本研究探讨提示语气对LLM相关性评判的影响,发现语气主要改变评判的宽松度而非判断力,且对校准一致性影响大于排名结果,提示语气是绝对标签场景下的效度威胁。

AI 中文摘要

大型语言模型越来越多地被用作相关性评判员,然而它们的标签可能会随着提示的表面形式而改变。我们研究了这样一个特征——语气——在3,498个TREC DL19/DL20查询-段落对上,跨越八个评判模型、五个分类器校准的礼貌级别以及每个级别的三个释义。效果强烈依赖于模型:一个评判员显示出结构化的U形响应,而大多数仅显示出微小变化。在语气改变一致性的地方,结果更符合评判员严重性操作点(其整体评分宽松度)的偏移,而非判断力的提升。一致性随此偏移将评判员移向或远离人类标注者的严格程度而上升或下降。查询不相交的交叉拟合保留了预期的关联(Spearman ρ = -0.683;精确模型块置换 p = 0.019)。语气对基于校准的一致性的影响大于对排名结果的影响:在32个模型-语气对比中,NDCG@10的最大绝对平均变化为0.011,尽管Kendall's τ低至0.743表明重排序减少但未消失。该解释调和了先前矛盾的发现,并指出当绝对相关性标签重要时,提示语气是一个潜在的效度威胁。

英文摘要

Large language models are increasingly used as relevance judges, yet their labels can shift with prompt surface form. We study one such feature -- tone -- on 3,498 TREC DL19/DL20 query-passage pairs, across eight judge models, five classifier-calibrated politeness levels, and three paraphrases per level. Effects are strongly model-dependent: one judge shows a structured U-shaped response, whereas most show only small changes. Where tone changes agreement, the results are more consistent with a shift in the judge's severity operating point -- its overall scoring leniency -- than with improved judgment. Agreement rises or falls as this shift moves the judge toward or away from human annotators' strictness. A query-disjoint cross-fit retains the expected association (Spearman $ρ= -0.683$; exact model-block permutation $p = 0.019$). Tone affects calibration-based agreement more than ranking outcomes: across 32 model-tone contrasts, the largest absolute mean change in NDCG@10 is 0.011, although Kendall's $τ$ as low as 0.743 shows that reordering is reduced, not absent. The account reconciles prior contradictory findings and identifies prompt tone as a potential validity threat when absolute relevance labels matter.

Comments3 pages, 2 figures, 2 tables. Accepted at the 20th ACM Conference on Recommender Systems (RecSys '26), Reproducibility and Practice Notes track. Code, prompt variants, and collection pipeline: https://github.com/dukesky/politeness-llm

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑