发表机构
University of California, Los Angeles; Shanghai Jiao Tong University; Arizona State University; Eastern Institute of Technology(加利福尼亚大学洛杉矶分校; 上海交通大学; 亚利桑那州立大学; 东方理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出技能优化中无参考评判门的预部署诊断方法,发现评判器能力需超1/k才具可用性,其基准准确率高估关键能力,可预测门控错误。
AI 中文摘要
文本空间技能优化通过演化自然语言技能文档来调整冻结智能体,每个候选方案需经验证门接受。现有验证门依赖可验证奖励,将这些方法限制在具备自动验证器的任务中。用LLM评判门替代验证器可突破该限制,但此类门是否携带可用信号尚未得到验证。我们提出一个前置问题:在将评判器纳入循环之前,能否判断其评分是否能区分正确与错误答案?我们将无参考评判器形式化为潜在求解器——其裁决基于与自身结论的一致性,因此其评估能力受限于其求解能力。该模型得出了评判器能力c和答案空间大小k的判别性(ROC-AUC)闭式边界,必要条件为c > 1/k,且边际AUC受项目难度混淆,而问题内估计量则不受影响。非介入式探测记录了评判器在真实优化运行中的评分,未改变任何决策。我们发现,当能力接近下限时分判别性处于随机水平,高于下限则具备可用性;评判器的基准准确率高估了关键能力;在闭环研究中,该筛选可预测发生何种类型的门控错误。研究成果是一种廉价的评判门部署前诊断工具。
英文摘要
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a judge in the loop, whether its scores separate correct from incorrect answers at all? We formalize a reference-free judge as a latent solver -- its verdict rests on agreement with whatever it would itself conclude, so its capacity to evaluate is bounded by its capacity to solve. The model yields a closed-form bound on discriminability (ROC-AUC) in the judge's competence $c$ and answer-space size $k$, a necessary condition $c > 1/k$, and the result that the marginal AUC is confounded by item difficulty while a within-question estimator is not. A non-intervening probe records judge scores on genuine optimization runs without altering any decision. We find discriminability at chance where competence sits near the floor and usable above it; that a judge's benchmark accuracy overstates the competence that matters; and, in a closed-loop study, that the screen predicts which kind of gating error occurs. The result is a cheap pre-deployment diagnostic for judge gates.
Comments9 pages, 4 figures, 5 tables