AI 中文总结
SkillTV-Bench是含681个案例的智能体轨迹基准,用于评估技能感知轨迹验证;提出SkillTV-Evolve优化JudgeSkill,使智能体评判者准确率提升14.8个百分点,选定轨迹成功率大幅提高。
AI 中文摘要
大型语言模型(LLM)智能体越来越多地通过工具使用和环境交互执行长周期任务,评估方式也从最终响应评分转向对完整执行过程的验证。对于技能增强型智能体,验证还需要任务时间技能中编码的程序性知识,因为这类知识指明了要检查的证据以及哪些失败是任务关键的。然而,现有的评判者基准通常仅暴露最终响应或静态轨迹,很少将任务时间技能与可直接检查的人工制品和环境结合起来。因此,我们推出SkillTV-Bench,这是一个包含681个案例的基准,来自11个领域的50个真实智能体轨迹,旨在评估面向LLM作为评判者(LLM-as-a-Judge)和智能体作为评判者(Agent-as-a-Judge)方法的技能感知轨迹验证。此外,我们提出SkillTV-Evolve,它将验证知识外部化为可复用的JudgeSkill,指导智能体评判者规划针对性检查并出具基于证据的裁决。在不相交的开发池中,自动化进化循环会利用判断错误的案例进一步优化JudgeSkill。在SkillTV-Bench上,优化后的技能使同一智能体评判者的准确率提高了14.8个百分点;在离线推出池选择中,它使选定轨迹的成功率从1次推出时的22.9%提升至10次推出时的45.5%。代码和数据可在此https URL获取。
英文摘要
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench