arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillTV-Bench:评估评判者在技能增强型智能体执行任务中的表现

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Zhi Han, Chenxi Zeng, Liuhaichen Yang, Zihan Guo, Ming Zhou, Yang Li

arXiv 2608.05573首次发表:更新:

AI 中文总结

SkillTV-Bench是含681个案例的智能体轨迹基准,用于评估技能感知轨迹验证;提出SkillTV-Evolve优化JudgeSkill,使智能体评判者准确率提升14.8个百分点,选定轨迹成功率大幅提高。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地通过工具使用和环境交互执行长周期任务,评估方式也从最终响应评分转向对完整执行过程的验证。对于技能增强型智能体,验证还需要任务时间技能中编码的程序性知识,因为这类知识指明了要检查的证据以及哪些失败是任务关键的。然而,现有的评判者基准通常仅暴露最终响应或静态轨迹,很少将任务时间技能与可直接检查的人工制品和环境结合起来。因此,我们推出SkillTV-Bench,这是一个包含681个案例的基准,来自11个领域的50个真实智能体轨迹,旨在评估面向LLM作为评判者(LLM-as-a-Judge)和智能体作为评判者(Agent-as-a-Judge)方法的技能感知轨迹验证。此外,我们提出SkillTV-Evolve,它将验证知识外部化为可复用的JudgeSkill,指导智能体评判者规划针对性检查并出具基于证据的裁决。在不相交的开发池中,自动化进化循环会利用判断错误的案例进一步优化JudgeSkill。在SkillTV-Bench上,优化后的技能使同一智能体评判者的准确率提高了14.8个百分点;在离线推出池选择中,它使选定轨迹的成功率从1次推出时的22.9%提升至10次推出时的45.5%。代码和数据可在此https URL获取。

英文摘要

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑