arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将交互历史转化为执行状态:面向长程编码智能体的运行时层

Turning Interaction History into Execution State: A Runtime Layer for Long-Horizon Coding Agents

Zehao Wang, Yisen Xu, Chenglin Li, Chao Peng, Bram Adams, Ahmed E. Hassan, Tse-Hsun, Chen

arXiv 2608.00808首次发表:更新:

AI 中文总结

该研究针对长程编码智能体易受过时状态干扰的问题,提出Ledger运行时层,通过显式执行状态提升编码任务性能并降低成本,效果在多个模型和基准上得到验证。

AI 中文摘要

长程编码智能体的轨迹中会积累数百个动作和观测结果,但该记录中没有任何信息表明哪些观测结果仍能描述仓库的当前状态。在每次决策前,模型必须从原始历史中隐式推断执行状态,当这种推断不足时,智能体会基于过时的文件内容采取行动,或重新执行结果仍有效的工作。我们提出Ledger,一个确定性的运行时层,它将智能体已完成的交互提炼为显式的执行状态:已观测的内容、已修改的内容以及已尝试的内容。Ledger将该状态存储在在线执行账本中,并在每一步的两个边界处应用。在模型行动前,信息路径会向提示中添加紧凑的运行时状态视图;在拟议命令运行前,管控路径会对照账本检查该命令,返回仍有效的早期结果以替代重新执行,并标记可能冗余的重复操作。该层不增加大语言模型调用次数,且包装了原本未修改的智能体。在全部500个SWE-bench Verified实例中,Ledger使用GPT-5 mini时将Pass@1从56.2%提升至64.2%,使用MiniMax M2.5时从75.8%提升至81.0%,同时总成本分别降低28.9%和31.8%;附加到OpenAI Codex时,其Pass@1提升3.4个百分点,成本降低24.4%。消融实验表明,大部分解决性能提升来自管控路径,大部分效率提升来自信息路径,二者结合表现最佳。我们得出结论,长程智能体缺乏的不是更短的历史视图,而是对自身执行状态的显式记录。

英文摘要

Long-horizon coding agents accumulate hundreds of actions and observations in their trajectories, yet nothing in this record indicates which observations still describe the repository as it currently stands. Before every decision, the model must implicitly infer the execution status from raw history, and when this inference falls short, the agent acts on outdated file contents or re-executes work whose results are still valid. We propose Ledger, a deterministic runtime layer that distills an agent's completed interactions into an explicit execution state: what has been observed, what has been modified, and what has been attempted. Ledger keeps this state in an online execution ledger and applies it at two boundaries of every step. Before the model acts, an inform path appends a compact runtime state view to the prompt; before a proposed command runs, a govern path checks it against the ledger, returning still-valid earlier results in place of re-execution and flagging likely-redundant repetition. The layer adds no language-model calls and wraps an otherwise unmodified agent. Across all 500 SWE-bench Verified instances, Ledger raises Pass@1 from 56.2% to 64.2% with GPT-5 mini and from 75.8% to 81.0% with MiniMax M2.5, while cutting total cost by 28.9% and 31.8%. Attached to OpenAI Codex, it adds 3.4 percentage points of Pass@1 at 24.4% lower cost. Ablations attribute most of the resolution gain to govern and most of the efficiency gain to inform, with their combination performing best. What long-horizon agents lack, we conclude, is not a shorter view of their history but an explicit account of their own execution state.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑