arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36319cs.AI

StateTape:面向长周期编码智能体的动作条件证据生命周期建模

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque

首次发表
浏览论文内容

中文总结 AI 辅助

针对长周期编码智能体上下文增长问题,提出StateTape框架,通过符号级代码图和写入磁带建模证据生命周期,清除失效记录并检索所需信息,在多个基准上提升解决率。

中文摘要 AI 辅助

尽管基于大型语言模型的编码智能体近期取得了成功,但在长周期任务中运行它们仍然具有挑战性,因为每次观察都会被附加到上下文中,导致上下文随每次观察而增长。基于历史的维护是一种常见的补救措施,它通过掩盖或总结旧观察,或修剪模型认为无用的内容,以较低的成本限制上下文长度。然而,这种方法仅从历史的文本内容进行决策,无法看到代码之间的连接关系。由于编码智能体在单个任务中会多次编辑代码,且每次写入都可能改变代码其他部分的含义,这种维护方式可能会保留已被写入操作证伪的记录,丢弃仍然有效的记录,并遗漏智能体下一步需要的代码。为克服这些挑战,本文提出StateTape,一种新颖且可扩展的框架,它根据仓库的变化而非上下文的增长来重写编码智能体的上下文。StateTape的核心思想是将仓库建模为符号级代码图,其依赖关系和语言规则揭示了哪些符号可能受到写入操作的影响。在此图上,磁带标记每次写入所改变的符号,从而将过时性从对文本的推断转变为对智能体写入操作的观察。我们提出了一种逐写入过程,其中磁带提名可能被写入操作证伪的记录,而一个小型管理模型则解决写入日志无法确定的问题,并进一步提供了理论分析和TraceBench基准,该基准将智能体持有的内容与实际所需内容进行标注对比。实验表明,StateTape能有效清除被证伪的记录并检索所需内容,从而在六个编码智能体和三个编辑密集型基准的所有实验中实现更高的解决率,且计算开销很小。

英文摘要

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent's context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent's writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.

发表机构

  • Causal Dynamics Lab(因果动力学实验室)
  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑