arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

停止与路由大语言模型(LLM)评审小组

Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao

arXiv 2608.19802首次发表:更新:

发表机构

School of Computer Science and Engineering; Sun Yat-sen University(计算机科学与工程学院; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将LLM评审小组设计建模为角色条件分配问题,提出策略优化评审调用,经多类审计对比,生成可复用的评审调用计划。

AI 中文摘要

LLM评估流程通常包含大量候选评审:通用LLM作为评审的提示词、奖励模型、安全分类器、置信度变体以及任务特定验证器。部署时的问题不仅在于哪种评审最佳,还应调用哪些评审、针对哪些样本,以及何时停止构建评审小组。我们将评审小组设计建模为角色条件分配问题。该方法从少量带标注的审计集、已声明切片和评审成本出发,估计目标相关角色:副本不添加条件信息、补充项提升全局评审效果、专家仅在特定切片上提供帮助。这些角色引出策略:丢弃副本、全局添加补充项、有条件路由专家,且当验证增益低于阈值时停止。在推理、代码、安全、偏好、奖励模型、摘要和数学审计中,该方法与单一评审、扁平评审小组、匹配多样性启发式、全调用堆叠、可靠性评审团及节俭级联进行对比。结果得到评审调用的机制图:在可部署切片上路由专家、在饱和验证器机制中停止、当风险收益值得成本时保留广泛集成、忽略条件副本。输出是可复用、可审计的下一批评估调用计划。

英文摘要

LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop. We formulate judge-panel design as a role-conditioned allocation problem. From a small labeled audit set, declared slices, and judge costs, the method estimates target-relative roles: copies add no conditional information, complements improve the global panel, and specialists help only on slices. These roles induce a policy: drop copies, add complements globally, route specialists conditionally, and stop when validation gain falls below a threshold. Across reasoning, code, safety, preference, reward-model, summarization, and math audits, the method is compared with single judges, flat panels, matched diversity heuristics, full-call stacking, reliability juries, and frugal cascades. The result is a regime map for judge calls: route specialists on deployable slices, stop in saturated verifier regimes, keep broad ensembles when their risk benefit is worth the cost, and ignore conditional copies. The output is a reusable, auditable call plan for the next evaluation batch.

Comments21 pages, 2 figures, 20 tables. Accepted at WISE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑