arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体技能轨迹的基于接地检查表的部分 credit 评估

Grounded Checklist Partial Credit for Agent Skill Trajectories

Suliu Qin, Lu Yin, Xilu Wang

arXiv 2608.27487首次发表:更新:

AI 中文总结

针对智能体轨迹评估的二元评分缺陷,提出 GCPC 方法,结合人工规则与 LLM 生成检查表,在 SkillsBench 等数据集上验证其评估区分度、与人类评估一致性及跨数据集迁移性更优。

AI 中文摘要

语言模型智能体越来越多地在交互环境中处理长程任务,但它们的评估通常依赖于任务级别的成功率,将整个执行轨迹简化为是否通过官方验证器的结果。这种二元评分隐藏了部分进展,对于程序性智能体技能评估尤其受限,因为技能可能在不改变最终结果的情况下改变执行过程。虽然检查表通过对单个任务要求评分提供了更细粒度的评估,但手动编写成本高且自动生成不可靠,使得可信评估难以规模化。为应对这些挑战,我们引入了基于接地检查表的部分 credit 评估(GCPC),这是一种由人工主导、大语言模型(LLM)实例化的智能体轨迹部分 credit 评估方法。人类只需定义一次可复用规则,LLM 即可基于任务指令和官方验证器实例化特定任务的接地检查表。为使判断与证据关联,评估者仅根据执行日志证据对每个项评分,证据缺失时弃权(不执行)。随后,单独的脚本步骤将官方验证器结果应用于评分。在包含 4455 条轨迹的去重 SkillsBench 评估集中,GCPC 在共享子集上区分官方通过(PASS)和失败(FAIL)结果的能力优于整体评估(AUC 为 0.689 对比 0.619)。对 12 个任务的 96 条轨迹进行的人工评估显示,GCPC 与人类对进展的评估更一致。应用于 1946 对匹配的有/无技能的轨迹时,GCPC 揭示了 pass@1 隐藏的效应:在 879 对二元结果未改变的轨迹中,20.9% 的评分提升超过 0.10,18.7% 的评分下降幅度相同。GCPC 流程还可迁移至 Terminal-Bench 和 SWE-bench,证明其适用于技能条件评估之外的场景。

英文摘要

Language-model agents increasingly tackle long-horizon tasks in interactive environments, yet their evaluation commonly relies on task-level success rates by reducing an entire execution trajectory to whether the task passes an official verifier. This binary score hides partial progress and is particularly limited for procedural agent skill evaluations, since a skill can alter execution without changing the final outcome. While checklists provide finer-grained evaluation by scoring individual task requirements, costly manual authoring and unreliable automatic generation make trustworthy evaluation difficult to scale. To address these challenges, we introduce Grounded Checklist Partial Credit (GCPC), a human-governed and LLM-instantiated partial-credit evaluation of agent trajectories. Humans define reusable rules once, from which an LLM instantiates a task-specific checklist grounded in the task instruction and official verifier. To keep judgment tied to evidence, a judge scores each item from execution log evidence alone and abstains when evidence is missing. A separate scripted step then applies the official verifier outcome to the score. Across a 4,455-trajectory, deduplicated SkillsBench evaluation population, GCPC better discriminates official PASS and FAIL outcomes than holistic judging on the shared subset (AUC 0.689 vs. 0.619). Human evaluation on 96 trajectories from 12 tasks shows that GCPC aligns more closely with human assessments of progress. Applied to 1,946 matched with/without-skill pairs, GCPC exposes the effects hidden by pass@1: among 879 pairs whose binary outcome does not change, 20.9% improve by more than 0.10 while 18.7% regress by the same margin. The GCPC pipeline also transfers to Terminal-Bench and SWE-bench, demonstrating applicability beyond skill-conditioned evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑