arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRACE:用于自动化上下文工程的轨迹归因

TRACE: TRajectory Attribution for Automated Context Engineering

Yikai Zhao, Pradeep Kumar Misra, Saurabh Pandey

arXiv 2608.09153首次发表:更新:

发表机构

Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TRACE是利用历史智能体轨迹挖掘上下文故障的自动化反馈循环,可在上下文层诊断修复超80%的生产AI智能体故障,归因率72.7%、修复有效性82%。

AI 中文摘要

生产环境中的AI智能体在其上下文来源(系统提示、知识库、工具描述及过程技能)存在错误或缺口时会失效。当前维护依赖人工日志审查与临时调试,随着交互量增长形成可扩展性瓶颈。本文提出TRACE(TRajectory Attribution for Automated Context Engineering),一种利用历史智能体轨迹挖掘上下文故障诊断与修复的自动化反馈循环。核心洞见为轨迹蕴含用户修正、重述、放弃提示等隐式不满信号,可精准揭示上下文来源的故障位置,无需显式收集反馈。与模型微调不同,TRACE在上下文层运行,无需重新训练即可快速迭代。本文贡献包括:(1)轨迹挖掘框架,系统从历史智能体执行中提取诊断信息;(2)多组件因果归因,将文本梯度从整体提示优化扩展至异构上下文来源(技能、知识库、工具、提示);(3)探索性验证,智能体主动读取上下文来源,区分需“创建”的内容缺口与需“更新”的陈旧内容,达到96%的操作准确率;(4)可复用模拟方法与可验证基准,解决上下文调试公开数据集缺失问题,含六类故障分类、真值标注及跨层验证协议。在覆盖三个复杂度层级(最多16个执行节点)的60条不满轨迹上,TRACE实现72.7%的根本原因归因率与82%的端到端修复有效性,表明生产系统中被忽视的历史轨迹资源可自动诊断并修复超80%的上下文层故障。

英文摘要

Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps. Current maintenance relies on manual log review and ad-hoc debugging, creating a scalability bottleneck as interaction volume grows. We present TRACE (TRajectory Attribution for Automated Context Engineering), an automated feedback loop that mines historical agent trajectories to diagnose and remediate context failures. Our key insight is that trajectories are rich with implicit dissatisfaction signals -- user corrections, rephrasing, abandonment cues -- that reveal precisely where context sources failed, without explicit feedback collection. Unlike model fine-tuning, TRACE operates on the context layer, enabling rapid iteration without retraining. We make four contributions: (1) a trajectory mining framework that systematically extracts diagnostic information from historical agent executions; (2) multi-component causal attribution that extends textual gradients from monolithic prompt optimization to heterogeneous context sources (skills, knowledge bases, tools, prompts); (3) exploratory verification, where agents actively read context sources to distinguish content gaps requiring CREATE from stale content requiring UPDATE, achieving 96% operation accuracy; and (4) a reusable simulation methodology and verifiable benchmark addressing the absence of open datasets for context debugging, with a six-category fault taxonomy, ground truth annotations, and a cross-layer verification protocol. On 60 dissatisfaction traces spanning three complexity tiers (up to 16 execution nodes), TRACE achieves 72.7% root cause attribution and 82% end-to-end fix effectiveness, showing that over 80% of context-layer failures can be automatically diagnosed and remediated by mining historical trajectories, an overlooked resource in production systems.

Comments21 pages, 5 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑