共享评判,学习弃权:专业化在大语言模型评估中的作用
Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation
浏览论文内容
中文总结 AI 辅助
该研究探讨大语言模型评估中领域专业化的架构设计,发现将专业化置于弃权而非评判环节,通过共享学习与级联路由可提升评估准确率并降低计算量。
中文摘要 AI 辅助
智能体系统扩大了生成候选输出与审查候选输出之间的差距。本文提出一个实用的架构问题:领域专业化应内置到评估器的权重中,还是内置到决定其判断何时可信的规则中?我们研究了99952个公开的、基于评分规则的示例。提供正确的评分规则比仅使用响应的对照组提高了2.11个百分点的锁定测试准确率;替换为不相关的评分规则则会损失2.66个百分点。然而,将相同的训练语料库分配给8个标准族LoRA评估器,会损失10.05个百分点,且在5%风险目标下,经审核的覆盖率从24.44%降至5.43%。将存储容量与一个秩为64的适配器匹配并未重现此损失,该结果也无法用学习率或优化器步骤解释。从共享的训练评估器初始化族适配器,可将测试准确率恢复至76.85%,比相同学习率下从头训练高出19.94个百分点(95%置信区间为18.88-21.02)。当专业化管控弃权(不执行)而非评判时,结果会发生变化。在RewardBench 2上,学习到的正确性头将示例路由通过0.6B-4B-8B级联,且不改变任何奖励分数。在20次锁定重分区中,该级联达到89.40%的准确率,而单独使用8B模型的准确率为84.75%,归一化参数计算量为0.415。所有运行均通过精确的单侧95%风险审核;基于边际的规则在使用至少0.94计算量时,准确率仍接近84.8%。这些结果提出了一条有条件的设计规则:共享评判的学习,直到有足够数据证明拆分合理为止,并将特定领域的适配置于经审核的发布边界中。
英文摘要
Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weights, or rule-based deferral policies for safe judgment acceptance. On 99,952 rubric-conditioned samples, correct rubrics improve accuracy by 2.11 points, while incorrect rubrics reduce performance by 2.66 points. Splitting training data across eight criterion-specific LoRA experts lowers accuracy by 10.05 points and reduces 5% error-bound coverage from 24.44% to 5.43%. This loss is independent of model size and training settings, with most performance recoverable by warm-starting experts from a unified judge. Sweeping data budgets confirms scratch expert specialization yields no empirical gains. Warm-started splits appear competitive with unified models, yet under limited data, unified training outperforms split training, with specialization beneficial only after unified training plateaus. Results hold on HealthBench, where physician rubrics improve accuracy while flawed rubrics degrade performance. Unlike weight specialization, deferral policies enable efficient evaluation. On RewardBench 2, lightweight deferral heads form a 0.6B-4B-8B reward cascade with no core scoring modification. Across 20 splits, the cascade achieves 89.40% accuracy versus 84.75% for a standalone 8B judge at 41.5% compute, satisfying 95% risk constraints. Margin-based deferral matches accuracy at far higher compute cost. The design generalizes across models, improving Tulu-3-8B and Skywork-8B performance with a lightweight DeBERTa frontend. We derive simple, robust evaluator design rules: unify judgment training or warm-start split models, and use audited deferral cascades for low-cost, reliable LLM evaluation.
发表机构
- Cheung Kong Graduate School of Business(长江商学院)
机构由 AI 辅助整理,请以论文原文为准。