编码智能体的自我修正应携带多少证据?用于自蒸馏的自适应狄利克雷证据
How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation
浏览论文内容
中文总结 AI 辅助
本文提出有效证据自蒸馏(EESD)方法,通过分离执行相关性与证据质量,利用狄利克雷后验为编码智能体的自我修正提供自适应学习权重,实验表明该方法显著降低NLL并提升Pass@1。
中文摘要 AI 辅助
执行反馈使编码智能体能够修改程序并从自身的修正中学习。修正的学习权重应同时反映其执行所支持的状态转移以及该支持背后的证据量。我们引入了有效证据自蒸馏(EESD),该方法分别表示这些量。归一化的执行相关性决定相对转移支持和有效伪计数质量;随后,狄利克雷后验产生一个不确定性惩罚权重,用于基于KL锚定的修正学习。在对称先验下,改变质量保持类别排序,且有效质量产生的监督系数受其匹配的固定质量对应物的约束。在四个模型-领域历史扫描中,将可见观测从1增加到8,将未来结果的负对数似然(NLL)降低了55.0%-59.3%。在8个观测时,有效质量在所有四次比较中均实现了比固定质量更低的NLL。在主要的匹配DeepSeek/RunBugRun研究中,argmax预测在全部3,000个示例上一致,且在集中相关性下获得了最大的NLL增益。经过一轮修正学习后,DeepSeek/CodeARC全测试Pass@1从15.0%提高到20.4%,配对95%源自助法区间为[+2.8, +8.0]个百分点。十二种设置的下游评估确定了该更新的模型-领域范围。这些结果表明,将证据支持与证据质量分离如何改变编码智能体中的概率估计和修正学习。
英文摘要
Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
发表机构
- University of Cambridge(剑桥大学)
- Western Sydney University(西悉尼大学)
- Guangdong University of Technology(广东工业大学)
机构由 AI 辅助整理,请以论文原文为准。