arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06909cs.AI

长时程智能体轨迹归因:统一基准与细粒度标注框架

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

Jing Chen, Yang Sun, Li Zhang, Lin Xu, Jie Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有智能体轨迹基准缺乏细粒度归因分析支持的问题,提出轨迹归因任务,构建含1300余条标注轨迹的统一基准与标注框架,定义两项评估任务并发布可复用标注技能。

中文摘要 AI 辅助

大型语言模型(LLM)智能体越来越多地通过包含用户指令、工具使用、外部观察和记忆的长时程轨迹运行。现有基准主要评估行为结果,但对细粒度归因分析的支持有限。我们引入轨迹归因概念,并为此任务开发了一个基准和标注框架。该基准在统一组件模式下组织异构轨迹,并提供主要归因组件的标注,以及适用情况下的攻击链和执行链。通过实例化来自AgentDojo的轨迹以及Agent3Sigma的Stage和Canary设置的轨迹,生成了超过1300条标注轨迹,涵盖任务对齐动作、不安全动作和安全弃权(不执行)。该基准定义了两个评估任务:主要归因定位和归因链恢复,并提供了基于增量轨迹贡献和组件级留一法扰动的参考基线。它捕获了多样化的归因设置,包括局部和长程归因以及结构化归因链。参考基线结果在这些设置中表现出显著的性能差异,为基准的归因挑战提供了初步特征。除了此初始实例化之外,我们还发布了可复用的标注技能,该技能可使新智能体模型生成的轨迹在同一框架下实现标准化、标注和评估。项目资源及未来版本可访问此https URL获取。

英文摘要

Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.

发表机构

  • Huawei Technologies Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑