arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Otap:用于评估智能体轨迹中规划与执行的结构感知最优传输

OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

Babak Barazandeh, Subhabrata Majumdar, George Michailidis

arXiv 2607.17082首次发表:更新:

AI 中文总结

研究智能体轨迹评估问题,将其重新定义为执行图与有效解决方案图的距离,通过不平衡融合Gromov-Wasserstein传输问题实例化得到\otap{}分数,该分数可区分有效与无效轨迹,对保持依赖的重新排序不变且对冗余步骤有界敏感。

AI 中文摘要

大型语言模型智能体通过生成交织规划、工具调用和中间结果的轨迹来解决任务。当前评估指标将此类轨迹简化为二元成功标志或通过精确匹配与参考进行比较。成功标志无法区分合理解决方案与靠运气成功的方案,也无法说明失败运行出错的原因。精确匹配会惩罚有效但与参考顺序不同或分解方式不同的计划。我们将轨迹评估重新定义为智能体执行图与一组有效解决方案图之间的距离,并通过属性依赖图上的不平衡融合Gromov-Wasserstein传输问题来实例化它。由此产生的分数,称为\otap{}(智能体规划的最优传输),是一种伪度量,可证明对保持依赖的重新排序不变,对冗余步骤具有有界敏感性。其不平衡边际处理缺失或幻觉步骤而不强制匹配,其软耦合适应计划粒度的变化。在受控扰动和三个公共基准上,\otap{}在仅语义度量得分低于随机水平的情况下将有效轨迹与无效轨迹区分开来。当精确恢复依赖图时,其准确性最高,仅在从自由文本轨迹启发式推断图时下降。

英文摘要

Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag, compare it against a reference by exact matching, or delegate judgment to another language model. A success flag cannot distinguish a sound solution from one that succeeds by luck, and says nothing about why a failed run went wrong. Exact matching penalizes plans that are valid but reordered or decomposed differently from the reference. We reframe trajectory evaluation as a distance between the agent's execution graph and a set of valid solution graphs, and instantiate it via an unbalanced fused Gromov-Wasserstein transport problem over attributed dependency graphs. The resulting score, termed OTAP (Optimal Transport for Agentic Planning), is a pseudo-metric that is provably invariant to dependency-preserving reorderings and has bounded sensitivity to redundant steps. Its unbalanced marginals handle missing or hallucinated steps without forcing a match, and its soft coupling accommodates variation in plan granularity. On controlled perturbations and three public benchmarks, OTAP separates valid from invalid trajectories in a regime where semantics-only metrics score below chance. Its advantage tracks the fidelity of the dependency graph: largest where edges follow from operator semantics, smallest where they are inferred from free text. Where a formal verifier exists, strict surface metrics predict validity better than OTAP does, which places OTAP in open-ended domains where no verifier is available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑