发表机构
Independent(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对编码智能体修复失败示例易留隐患的问题,提出针对反例补充草图的智能合成方法,通过人类给出草图、智能体生成实现,依据反例修正并保留代码,经实验验证该方法能减少返工,承载审查策略。
AI 中文摘要
编码智能体可能在修复失败示例时未保留导致失败的领域规则,致使后续出现同样的错误。我们提出针对反例补充草图的智能合成方法,这是一种适用于在实现过程中发现管理策略的系统的存储库原生方法。人类先给出部分代码形式的草图,编码智能体生成首个实现。当具体失败暴露策略缺失或错误时,操作员明确批准修正后的行为和规则。智能体随后修订草图并针对该反例修复或重新生成代码及提示表面。完整存档保留出处;选定的回归集在揭示下一个候选者之前对每个修订进行把关;定期的完全重新生成测试进化后的草图而非提示历史或累积示例是否承载所学策略。我们用合成浏览器应用程序CatSynth及捕获的编码智能体实验演示了该方法。在使用GPT - 5.4 - mini的一次开放世界运行中,14个冻结候选案例中有8个成为反例。重建控制继承了该提升计划,所有三条路径都通过了8个已接受案例。从进化后的草图重建通过了21个保留案例中的19个,而从初始草图重建并重放所有已接受示例时通过了21个中的15个。跨反例保留代码需要9次开发者调用和719行累积工件变动,而重放所有案例需要15次调用和2394行,通过了21个保留案例中的18个。这些结果提供了可检查的证据,表明进化后的草图承载了经过审查的策略,并且在此次运行中保留代码减少了返工;但对于一个模型和一种揭示顺序,它们并未确立超出编码检查的普遍优越性或正确性。
英文摘要
Coding agents can fix a failing example without preserving the domain rule that made it fail. We present agentic synthesis against counterexample-supplemented sketches, a repository-native method for systems whose policy is discovered during implementation. A human starts with a partial sketch, and a coding agent compiles a replaceable projection. When simulation exposes missing or mistaken policy, an operator approves the corrected behavior and the minimum general rule the case authorizes. Every Developer call names its change authority and the rules, holes, anchors, and approved behavior that must survive. Conflict or ambiguous permission leaves the files unchanged and produces a clarification question. A complete archive preserves provenance; a curated regression set gates distinct boundaries. Before another candidate is revealed, the active case and curated regressions must pass both deterministic approved-output comparison and a separate review against the current sketch. Periodic clean regeneration tests whether the sketch carries the learned policy. We demonstrate the method with CatSynth, a captured synthetic application. In one open-world run with GPT-5.4-mini, 8 of 14 frozen candidates became counterexamples. Under the corrected protocol, replay-all, evolved-sketch rebuild, and retained Sketch-CE each passed all 8 accepted cases. They passed 14, 17, and 16 of 21 withheld cases, respectively. Sketch review rejected premature empty-input and tag policies and restored dropped anchors; adjudicated reviewer errors did not become policy. One model and one reveal order cannot establish general correctness or superiority. On this suite, the second check exposed drift hidden by deterministic replay, and the reviewed sketch passed three more withheld cases than raw example replay.
Comments32 pages, 5 displayed figures (4 distinct screenshots). Includes the CatSynth artifact supplement. Code and captured experiment artifacts: https://github.com/open-horizon-labs/counterexample-supplemented-sketches Clarifies the two-check CESS method and Developer change authority; adds the protocol-correct CatSynth rerun and replaces the prior withheld-case headline