面向开放文本多维度评估的维度间依赖关系
Inter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended Text
浏览论文内容
中文总结 AI 辅助
本文针对开放文本多维度评估中LLM评判者的维度间依赖问题,提出CorrGap度量该依赖关系,进而提出DimCheck方法缓解该问题,在多任务多模型上优于基线且推理成本低。
中文摘要 AI 辅助
LLM作为评判者的方法被广泛用于评估生成开放文本的质量。这类评估通常是多维度的,因为不同维度的文本错误模式可能存在差异,因此可靠的LLM评判者应独立评估每个目标维度。为量化LLM评判者在评估目标维度时对非目标维度的依赖程度(即维度间依赖关系),本文提出CorrGap方法。CorrGap通过比较不同文本组中LLM预测分数与真实分数的相关性差异来度量该依赖关系,研究发现维度间依赖关系在开放文本评估任务的LLM评判者中普遍存在。为缓解维度间依赖关系,本文提出DimCheck方法,该方法以逐步迭代的方式从LLM评判者生成的思维链(COT)中移除无关证据。实验结果表明,DimCheck可缓解维度间依赖关系,在三个LLM和四项任务上均优于强基线方法;此外,规模较小的训练后LLM可在DimCheck中近似规模更大的LLM,且推理成本显著更低。
英文摘要
LLM-as-a-judge methods are widely used for evaluating the quality of generated open-ended text. Such evaluations are generally multi-dimensional, since the error patterns in texts can be different for different dimensions. Therefore, reliable LLM judges should evaluate each target dimension independently. To quantify the extent to which LLM judges depend on non-target dimensions when evaluating a target dimension, i.e., inter-dimension dependence, we propose CorrGap. To measure this, CorrGap uses the difference in correlations between LLM-predicted scores and ground truth scores across different groups of texts. Using CorrGap, we show that inter-dimension dependence is pervasive across LLM judges in open-ended text evaluation tasks. To mitigate inter-dimension dependence, we propose DimCheck, a method that iteratively removes unrelated evidence from COTs generated by LLM judges in a step-wise way. We show that DimCheck mitigates inter-dimension dependence and outperforms strong baselines across three LLMs and four tasks. We also show that smaller trained LLMs can approximate larger LLMs in DimCheck, with much lower inference costs.
发表机构
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。