发表机构
Peking University; Weixin Al, Tencent Inc(北京大学; 腾讯公司微信智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM评审员多评分标准评估时的干扰问题,提出SARA方法,通过策略内自蒸馏提升评估一致性,且该一致性可跨数据集迁移。
AI 中文摘要
大语言模型(LLM)评审员越来越多地依据细粒度评分标准清单对响应进行评估。当一个样本需要对应多个评分标准时,现有方法通常会在单独的推理调用中评估每个评分标准。在一次推理中评估所有评分标准是一种更高效的自然替代方案,但我们发现这会引入评分标准干扰:一个评分标准的判定结果会因其他评分标准的共存而发生改变。在一项初步研究中,当在不同组成的评分标准集下进行评估时,仅三分之一的样本获得了完全一致的判定结果。我们开发了一种测量框架,通过四项受控操作来探测干扰:评分标准集扩展、子集划分、重排序和噪声注入。为在无需外部监督的情况下缓解干扰,我们提出了自锚定评分标准对齐(Self-Anchored Rubric Alignment,SARA)方法。SARA利用模型自身的单评分标准判定作为稳定锚点,并通过策略内自蒸馏将多评分标准推理与这些锚点对齐。我们在三个数据集(HealthBench、FLASK、ResearchQA)和两个模型系列(Qwen3、Llama-3.1)上对SARA进行了验证。SARA在保持与基础模型以及作为参考评审员的GPT-4.1一致性的同时,持续提升评估一致性。此外,学习到的一致性可跨数据集迁移,证实SARA传授的是一种通用能力,而非拟合数据集特定模式。
英文摘要
LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. Evaluating all rubrics in a single pass is a natural alternative with greater efficiency, but we find that it introduces rubric interference: the verdict on one rubric shifts depending on which other rubrics are co-present. In a preliminary study, only one-third of samples receive fully consistent verdicts when evaluated under rubric sets of varying composition. We develop a measurement framework that probes interference through four controlled operations: rubric set expansion, subsetting, reordering, and noise injection. To mitigate interference without external supervision, we propose Self-Anchored Rubric Alignment (SARA). SARA uses a model's own single-rubric judgments as stable anchors and aligns multi-rubric reasoning with these anchors through on-policy self-distillation. We validate SARA on three datasets (HealthBench, FLASK, ResearchQA) and two model families (Qwen3, Llama-3.1). SARA consistently improves evaluation consistency while maintaining agreement with both base models and GPT-4.1 as a reference judge. Furthermore, the learned consistency transfers across datasets, confirming that SARA teaches a general capability rather than fitting dataset-specific patterns.