发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的编码代理重复犯错问题,提出闭环框架,将审查评论编码为行为规则,经实验验证其能转移审查重点、降低错误复发率且跨接口转移,实现跨会话学习且不更新权重,积累人类工程智慧。
AI 中文摘要
基于大语言模型的编码代理在不同会话中会重复犯相同类型的错误,因为它们缺乏保留人类审查反馈修正的机制。我们提出了一个闭环框架,其中每个被接受的审查评论都被编码为一个持久的行为规则,逐步扩大代理可以自我检测的错误类集合。该框架将累积的规则集整合到一个版本控制的指令文件中,在代码提交前执行自我审查清单,并进行自动验证以确保规则集在增长时的完整性。在一个35多个服务微服务平台上进行部署时,规则集从5个行为规则、15多个特定语言标准和一个15项自我审查清单增长而来,所有这些都来自实际审查反馈。我们展示了11个记录的工作会话的实证结果,涵盖代码生成、拉取请求审查、事件调查和跨服务重构。我们观察到,累积的规则将审查工作从低级正确性转向设计级验证,实现了针对被裁定错误类别的0%复发率,并能跨异构代理接口转移。我们将我们的方法与经验性大语言模型学习(Reflexion、ExpeL、Voyager)和自动代码审查(CodeReviewer、SWE-bench代理)中的相关工作进行了比较,表明我们的框架在不更新权重的情况下实现了持久的跨会话学习,在生产代码库上运行而非合成基准,并解决了现有基准未测量的正交维度(随时间的行为一致性)。结果是一个编码代理,它在每个审查周期中都能改进,积累其人类合作者的工程智慧而不改变单个模型权重。
英文摘要
LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feedback. We present a closed-loop framework in which every accepted review comment is codified as a persistent behavioral rule, progressively expanding the set of error classes the agent can self-detect. The framework combines an accumulating rule set in a version-controlled instruction file, a self-review checklist executed before code submission, and automated validation that ensures rule set integrity as it grows. In deployment across a 35+ service microservices platform, the rule set grew from 5 to 18 behavioral rules, 15+ language-specific standards, and a 15-item self-review checklist, all derived from real review feedback. We present empirical results from 11 recorded working sessions spanning code generation, PR review, incident investigation, and cross service refactoring. We observe that accumulated rules shift review effort from low-level correctness toward design-level validation, achieve a measured 0% recurrence rate for ruled-against error classes, and transfer across heterogeneous agent interfaces. We compare our approach against related work in experiential LLM learning (Reflexion, ExpeL, Voyager) and automated code review (CodeReviewer, SWE-bench agents), showing that our framework achieves persistent cross-session learning without weight updates, operates on production codebases rather than synthetic benchmarks, and addresses an orthogonal dimension (behavioral consistency over time) that existing benchmarks do not measure. The result is a coding agent that improves with every review cycle, accumulating the engineering wisdom of its human collaborators without changing a single model weight.
CommentsAlready presented and accepted in - 32nd ICE IEEE/ITMC Conference (ICE 2026)