基于干预迁移的评分规则生成评估
Evaluating Rubric Generation with Interventional Transfer
浏览论文内容
中文总结 AI 辅助
本文提出干预迁移(IT)方法评估LLM生成的评分规则,通过HealthBench案例发现不同模型生成评分规则存在不对称性,该发现对LLM生成评分规则的应用有启示意义。
中文摘要 AI 辅助
针对AI基准测试中需要可靠评估且依赖特定专业知识的场景,实例特定评分规则十分常见,但该方法难以规模化,因此研究人员探索用大语言模型(LLM)生成评分规则。然而,即便有专家评分规则作为参考,如何规模化评估生成评分规则的质量仍不明确。本文提出一种名为干预迁移(Interventional Transfer, IT)的评分规则生成评估方法,其核心思想是:当对某一响应进行扰动使其通过或不通过某一评分规则时,若两个评分规则的表现同步变化,则二者相似。与现有评分规则生成评估方法不同,本文认为不同形式的干预迁移可用于评估生成评分规则在不同任务中的效用。在HealthBench案例研究中,本文将该方法应用于评估Qwen3.8-27B、Deepseek-V4-Flash、Opus-5生成的评分规则对GPT-5.6-Terra响应的评估,结果显示存在不对称性:根据生成/专家评分规则降低响应质量的扰动,会在对应的专家/生成评分规则中转化为更低分数;但在某一评分规则中提升响应质量的扰动,无法可靠地在另一评分规则中转化为更高分数。本文认为该发现对LLM生成评分规则在性能监控和 hill-climbing( hill-climbing 为保留的英文术语)中的应用具有启示意义,且本文方法与现有评分规则评估方法不同,未呈现本文观察到的这种不对称性。
英文摘要
Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.
发表机构
- Johns Hopkins(约翰斯·霍普金斯大学)
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。