arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VERA:面向自适应LLM评判器的判定条件可靠性

VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges

Qiushui Xu, Syamil Mohd Razak, Tao Yuan, Piotr Habas

arXiv 2610.05452首次发表:更新:

发表机构

Penn State University; Amazon(宾夕法尼亚州立大学; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VERA,利用隐藏激活按判定组估计可靠性,引导周期性适应框架,在多个基准上显著提升LLM评判器性能与焦点类别召回率。

AI 中文摘要

准确估计评判可靠性是将LLM评判器适应于新验证反馈同时保留先前学习行为的一个核心挑战。然而,现有方法通常依赖输出级置信度,这可能过于自信且与评判正确性对齐不佳。我们提出VERA,一种判定条件可靠性轴,通过在每个预测判定组内区分正确与不正确的判定,从隐藏激活中估计可靠性。使用VERA作为控制信号,我们开发了一个VERA引导的周期性适应框架,该框架整合了可靠性排序的纠正更新、可靠性残差回放以及可靠性方向的周期性刷新。在Chatbot Arena上进行VERA引导适应后,8B和14B参数评判器在四个保留的公开基准上均优于最强基线,相对提升高达23.01%。该框架还在一个独立的专有时间性审计任务上,相对于最强自适应基线,将焦点类别召回率相对提升了高达16.1%。

英文摘要

Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose VERA, a VErdict-conditioned Reliability Axis that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using VERA as a control signal, we develop a VERA-guided periodic adaptation framework that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to 23.01%. The framework also improves focal-class recall by up to 16.1% relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑