arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当噪声制造偏见:噪声文本下LLM-as-a-Judge偏见测量的脆弱性

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak

arXiv 2609.11067首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现噪声文本会使LLM评判者系统性高估社会偏见,且中性转有偏的概率远高于反向,最脆弱模型在轻度噪声下失真最严重。

AI 中文摘要

大型语言模型越来越多地被用作评判者来测量文本中的社会偏见,然而它们所评判的文本往往带有噪声,包含拼写错误、非正式拼写和破碎的标点符号。这种表面噪声对社会偏见测量的后果仍不清楚。为了探究这一问题,我们在多种强度水平下对3,822条与刻板印象相关的回复应用了五种现实的噪声条件,并将所得的偏见评判与对原始文本的评判进行比较。我们发现,这种表面噪声并不会对称地降低偏见测量的质量:它更有可能将中性评判转变为有偏评判,而非将有偏评判转变为中性评判,差距高达120倍。我们进一步在四个LLM评判者中观察到两个非显而易见的效果:在最脆弱的评判者中,失真在轻微、现实的噪声水平下最为纯粹,此时擦除最为稀少,并且随着评判者变得稳健,失真趋向于均等而非反转。因此,在噪声文本上测量的偏见被系统性高估,尤其是在对公平性最为重要的类别中。

英文摘要

Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.

Comments15 pages, 4 figures. Accepted at W-NUT 2026. Code: https://github.com/dong4918-skku/Fable

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑