Chronos 使代码智能体能够对软件演进进行推理
Chronos Enables Code Agents to Reason over Software Evolution
- Zhejiang University(浙江大学)
- Shanghai Jiao Tong University(上海交通大学)
- Hithink Research(航信研究院)
- National Industrial Information Security Development Research Center(国家工业信息安全发展研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出 Chronos 框架,通过关联历史提交经验优化代码智能体,在 SWE-Bench 等基准测试中提升了代码智能体的补丁解决率,验证了提交关系对代码智能体的价值。
AI中文摘要:
历史提交记录(pull requests)记录了代码库当前状态背后的设计决策、兼容性约束和实现模式。与新任务相关的经验可能涵盖描述不同关注点的相关变更。我们引入 Chronos,这是一个测试时框架,可为基于大语言模型(LLM)的代码智能体提供这种关联的历史信息。Chronos 将已合并的提交提炼为结构化的经验卡片,并通过代码级、开发者意图级和组织级关系的类型化图将它们关联起来。语义搜索识别入口卡片,加权多跳扩展检索相关变更以供选择性阅读。该记忆同时指导候选生成和补丁选择:一个专注于补丁的变更智能体和一个验证策略智能体各自生成一个补丁,而演进管理者则参考历史在两者之间进行选择。在 SWE-Bench Verified 上,完整工作流程在所有六个评估的 LLM 主干模型上均提升了 SWE-Agent 的性能,将平均解决率从 69.2% 提升至 72.9%,使用 MiniMax M2.5 时达到 79.8%。在相同主干模型下,其将 SWE-Bench Pro 的解决率从 48.3% 提升至 51.7%,将 FEA-Bench Lite 的解决率从 41.0% 提升至 43.5%。两种受经验指导的单智能体变体也均优于基础智能体。在针对 100 个任务、每个任务检索 10 张卡片的人工评估中,基于图的检索相比平面语义检索将有用卡片的平均数量从 1.24 提升至 2.87。这些结果证明了提交关系在检索有用仓库经验方面的价值,以及所评估的工作流程在补丁生成和选择过程中应用该经验的价值。
英文摘要:
Historical pull requests record the design decisions, compatibility constraints, and implementation patterns behind a codebase's current state. Experience relevant to a new task can span related changes whose descriptions emphasize different concerns. We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents. Chronos distills merged pull requests into structured experience cards and connects them through a typed graph of code-level, developer-intent, and organizational relations. Semantic search identifies entry cards, and weighted multi-hop expansion retrieves connected changes for selective reading. The same memory guides candidate generation and patch selection: a patch-focused change agent and a validation-strategy agent each develop a patch, and an evolution steward consults history to select between them. On SWE-Bench Verified, the full workflow improves SWE-Agent across all six evaluated LLM backbones, raising the mean resolution rate from 69.2% to 72.9% and reaching 79.8% with MiniMax M2.5. With the same backbone, it raises resolution rates from 48.3% to 51.7% on SWE-Bench Pro and from 41.0% to 43.5% on FEA-Bench Lite. Both experience-guided single-agent variants also outperform the base agent. In a human evaluation on 100 tasks with ten cards retrieved per task, graph-grounded retrieval increases the mean number of useful cards from 1.24 to 2.87 over flat semantic retrieval. These results demonstrate the value of PR relations for retrieving useful repository experience and of the evaluated workflows for applying that experience during patch generation and selection.