发表机构
University of Chinese Academy of Sciences; Pusan National University; Shenzhen University of Advanced Technology(中国科学院大学; 釜山国立大学; 深圳理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体评估,提出任务级自助认证方法,确保自动评判错误率低于预算,实现可认证的选择性自动化,并兼作自训练过滤器以零标签进入新领域。
AI 中文摘要
评估LLM智能体仍然以人工阅读轨迹结束,因为自动评判器无法保证其出错的频率。我们提出一个操作性问题:评判器可以接管多大比例的智能体评估,并附带一个认证,即自动判定轨迹中的错误率保持在预算alpha以下?智能体语料库抵制标准答案:许多智能体尝试相同的任务,因此轨迹以相关簇的形式到达,现有选择性评判方法的i.i.d.认证可能高估安全性:一个朴素的认证可以声称98%的自动化,而其实际错误率在17.5%的任务重采样中超过预算。我们引入一种任务级自助认证,在我们测试的每种情况下都有效,同时匹配朴素认证的覆盖率;有限样本的簇有效替代方案在现实任务数量下无法认证任何内容。在此认证下,一个使用SFT和拒绝加权GRPO训练的4B对数概率评判器在alpha=0.1时,对工具使用和网页语料库认证了0.30-0.59的评估,这是强引导的前沿模型中唯一在两个主要语料库上均通过认证的评判器。认证覆盖率在训练前仅凭基率和判别力即可预测(留一语料库R^2=0.96)。最后,该认证兼作自训练过滤器:在认证区域内收获的伪标签,其污染度按构造受alpha限制(六次收获中实际为0.000-0.041),使评判器在零目标训练标签的情况下以域内强度进入未见过的领域。
英文摘要
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.