arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布转移下可靠长时代理上下文演化的范围验证

Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift

Dan C. Hsu, Luke Lu

arXiv 2607.09175首次发表:更新:

发表机构

RedMind Research; National Taiwan University(红芯研究公司; 国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究长时代理上下文演化在分布转移下的验证问题,提出图正则化代理上下文演化方法(GRACE),通过维护指令为类型化语义图并在局部验证更新,实验表明该方法显著提升了严格可靠性。

AI 中文摘要

部署的语言模型代理依赖于代理上下文,即由操作框架组装的模型外部文本控制内容。在这项工作中,上下文的可变部分是一个持久的系统级指令,它在模型、工具和框架保持不变的情况下根据操作经验进行更新。在长时演化过程中,随着累积指令的增加和相互作用,纯文本维护使得验证变得越来越困难。我们提出了图正则化代理上下文演化(GRACE),它将持久指令组件维护为类型化语义图,并在修改节点的局部类型化邻域内验证提议的更新。接受的图更新被重建为对部署时使用的文本指令检查点的增量编辑。我们在从$\tau^2$-bench派生的固定电信代理框架内,在受控分布转移协议下评估GRACE。在五个独立复制中,GRACE将以pass^3衡量的严格可靠性从Gemini 2.5 Flash的零样本值0.091提高到最终检查点的0.673±0.136。这超过了在相同保留集上Gemini 3.1 Pro的零样本参考值0.242,而纯文本HCE基线最终为0.191±0.051。这些结果确定了可靠长时上下文演化的两个要求,一个使验证局部化的结构基础和一个使累积指令内容可用的整合机制。

英文摘要

Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, tools, and harness remain fixed. Over long evolution horizons, flat-text maintenance makes verification increasingly difficult as accumulated instructions grow and interact. We propose Graph-Regularized Agentic Context Evolution (GRACE), which maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. Accepted graph updates are reconstructed as incremental edits to the textual instruction checkpoint used at deployment. We evaluate GRACE within a fixed telecom agent harness derived from $τ^2$-bench under a controlled distribution-shift protocol. Across five independent replications, GRACE improves strict reliability, measured by pass^3, from the Gemini 2.5 Flash zero-shot value of 0.091 to 0.673$\pm$0.136 at the final checkpoint. This exceeds a Gemini 3.1 Pro zero-shot reference of 0.242 on the same held-out set, while the flat-text HCE baseline finishes at 0.191$\pm$0.051. These results identify two requirements for reliable long-horizon context evolution, a structural substrate that makes verification local and a consolidation mechanism that keeps accumulated instruction content usable.

Comments18 pages, 3 figs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑