arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceDance:一种从真实世界智能体部署轨迹构建智能体行为基准的自动化系统

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang, Ziyi Chen, Yan Zhang, Qinbo Bai, Mengyuan Chao, Jing Ning, Qiyue Hua, Huiyi Chen, Hanrong Zhang, Henry Peng Zou, Jie Yang, Wei Xu, Philip S. Yu

arXiv 2609.33295首次发表:更新:

AI 中文总结

TraceDance从真实部署轨迹自动构建针对性智能体行为基准,采用Anchor-and-Confirm与合成循环高效生成,实验显示前沿LLM平均通过率仅26.7%,助力递归自我改进。

AI 中文摘要

智能体在完成任务的过程中可能会表现出不良行为。开发者需要针对部署中遇到的具体行为进行测试,而不仅仅是固定的基准测试套件。我们提出了TraceDance,一个智能体系统,它能够从部署轨迹中针对用户指定的不良行为构建有针对性的基准。为了高效构建,Anchor-and-Confirm方法将可编程检索与Flash大型语言模型(LLM)的候选级确认相结合,而Anchor合成循环则生成并修订自定义行为的规范。这些基准采用决策点延续方法,在记录的决策点处评估LLM的下一轮行为,并使用特定于行为的评分标准,无需参考答案或环境重放。在编码和通用工具使用方面的实验利用了252,557个会话,生成了107个基准,包含4,125个实例,满足了95.3%的构建目标请求。人工标注者在84%的抽样实例中确认了所请求的行为,而自动评分器与人工通过/失败判断的一致性程度与标注者之间的一致性相当。九个前沿LLM的平均通过率仅为26.7%,表明它们在评估的决策点上仍然难以做出适当响应。跨行为特定基准的分析进一步揭示了当前LLM作为智能体行为方式的弱点。通过将部署问题转化为有针对性的基准,TraceDance可以作为递归自我改进(RSI)循环的关键组成部分。

英文摘要

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.

Comments34 pages, 7 figures. Project website: https://zhishanq.github.io/TraceDance/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑