发表机构
Japan Advanced Institute of Science and Technology(日本先进科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM-as-a-Judge的评分偏差问题,提出通过随机数生成识别并校正LLM潜在数字偏差的方法,在四项任务上的表现优于基线。
AI 中文摘要
大语言模型(LLM)常被用作文本质量的评估者,即LLM-as-a-Judge,其性能优于依赖参考文本的传统自动评估指标。但LLM评估者倾向于不考虑被评估文本的上下文而生成特定分数,这一现象被称为评分偏差。本研究提出一种缓解该评分偏差的新方法:指示LLM随机生成数字 token,通过测量观测到的数字分布与均匀分布的偏差来识别LLM的潜在数字偏差;在随机数生成的提示中添加使用LLM评估者的下游任务定义,以测量特定任务的潜在数字偏差;在LLM评估时,结合LLM的潜在数字偏差对给定输入的token生成概率进行校正。在四个不同任务(LLM对齐评估、摘要评估、语义文本相似度、语义文本相关性)上的实验结果表明,所提方法的性能优于基线方法,包括未去偏的LLM及先前的校准方法。此外,研究还证实评分偏差会因LLM、任务和分数范围而异,这表明根据具体情况测量潜在数字偏差具有重要意义。
英文摘要
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of the evaluated text, which is known as scoring bias. This study proposes a novel method to mitigate this scoring bias. An LLM is instructed to randomly generate number tokens, and the latent numerical bias of the LLM is identified by measuring the deviation of the observed distribution of numbers from the uniform distribution. A definition of a downstream task, for which an LLM evaluator is used, is added to the prompts for random number generation to measure task-specific latent number bias. In the evaluation by an LLM, the token generation probabilities for a given input are rectified considering the LLM's latent number bias. Results of the experiment on four different tasks, evaluation of LLM alignment, evaluation of summarization, Semantic Textual Similarity, and Semantic Textual Relatedness, demonstrate that our proposed method outperforms the baselines, including an LLM without debiasing and previous calibration methods. In addition, it is confirmed that scoring bias varies across LLMs, tasks, and score ranges, indicating the importance of measuring latent number bias as the case may be.