RADAR:用于LLM作为评判者评估的基于 Rubric 的依赖与冗余分析
RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation
浏览论文内容
中文总结 AI 辅助
本文提出 RADAR 框架,用于在 LLM 作为评判者的大规模评估前,通过少量探针估计评估标准间的耦合,验证显示其可恢复人类标准间高相关性,提供审计信号。
中文摘要 AI 辅助
基于 rubric 的 LLM 作为评判者的评估流程通常假设评估标准提供独立信号,但实际上标准在行为上可能存在耦合:改进某一标准会系统性改变另一标准的得分,进而影响模型发布或产品更新决策所用的总得分。本文提出 RADAR,一种轻量级预飞行诊断框架,用于在大规模评估前估计此类耦合。给定 rubric,RADAR 生成针对性合成探针,对所有标准评分,生成定向耦合矩阵以展示哪些标准共同评分及耦合方式。在 NVIDIA HelpSteer2、SumPubMed 和 Yale-Salesforce SummEval 三个行业相关评估场景中验证 RADAR,仅用每个标准少量探针,RADAR 可恢复人类的标准间相关结构(Pearson r > 0.84),为从业者提供关于冗余、层级和聚合敏感性的具体审计信号,无需开展大规模评判即可使用。
英文摘要
Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one criterion may systematically change scores on another, distorting aggregate scores used in model-release or product-update decisions. We introduce RADAR, a lightweight preflight diagnostic framework for estimating such coupling before large-scale evaluation. Given a rubric, RADAR generates targeted synthetic probes, scores each probe on all criteria, and produces a directional coupling matrix that shows which criteria co-score and how. We validate RADAR on three industry-relevant evaluation settings: NVIDIA HelpSteer2, SumPubMed, and the Yale-Salesforce SummEval benchmark. Using only a small number of probes per criterion, RADAR recovers human inter-criterion correlation structure (Pearson r > 0.84) and provides practitioners with concrete audit signals about redundancy, hierarchy, and aggregation sensitivity before committing to large-scale judging.
发表机构
- Microsoft(微软公司)
- University of Florida(佛罗里达大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。