arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

骨干自适应证据路由用于鲁棒的成对LLM评判

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

Zeyan Li, Jing Peng, Jianfeng Xu

arXiv 2609.30751首次发表:更新:

发表机构

Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出骨干自适应证据路由(BAER),通过自适应选择证据机制并在保持候选对称性的同时分离偏好与可靠性,在四个基准和两个8B骨干上均取得最高准确率,优于固定协议。

AI 中文摘要

成对语言模型评判器可以通过直接比较、推理或基于参考的验证来收集证据,但没有任何单一协议能在所有基准和评判骨干上表现最佳。我们提出了骨干自适应证据路由(BAER),它在保持候选对称性的同时自适应地调整证据机制:交换两个响应可能会反转偏好,但不能改变其强度。BAER将每个专家的符号化偏好与候选无关的可靠性分离,并构建了三个对称头:证据堆叠、基于可靠性的专家路由和候选盲参考验证。开发数据为每个基准-骨干条件选择一个头,该选择在测试前被冻结。在四个基准和两个8B评判骨干上,BAER在所有八个条件下均达到所比较方法中的最高测试准确率,具有完全的预测覆盖率,并且相对于最强外部基线提升了0.87-7.32个百分点。结果表明,自适应地调整证据收集方式比在所有地方固定一种评判协议更可靠。

英文摘要

Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliability-based expert routing, and candidate-blind reference verification. Development data select one head for each benchmark--backbone condition, and that choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER achieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0.87--7.32 points over the strongest external baseline. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑