arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从实验到决策:在自主编码研究中重用证据

From Experiments to Decisions: Reusing Evidence in Autonomous Coding Research

Bobber Cheng

arXiv 2609.13299首次发表:更新:

AI 中文总结

本研究通过分析自主编码智能体在实验证据重用中的缺陷,提出可检查的交接机制和晋升谓词,以保留实验结果及其对后续行动的限制。

AI 中文摘要

自主编码智能体能够记住一个实验,却可能沿用该实验并未证实的结论。我们重构了在一个包含400个任务的NeuroGolf活动中证据是如何被重用的,并选取了同一运营商的部分井眼预测记录作为跨领域对比。一个数值反例揭示了一种过度宽泛的排除;其他案例则区分了失败、不完整和修订后的程序分别证明了下一步该做什么。井眼记录揭示了另一个弱点:协调者正确承认了仅基于新颖性的拒绝,随后却将其总结为已测量的闭合。另外,在保留的实验中,基于前缀的接受门控在三个阈值下产生的目标分数均低于未门控的自适应。这些案例促使我们提出一种可检查的交接机制,该机制关联了被测试的命题、其适用范围、候选者与评估者身份、检查状态以及重新开启条件。一个调度树将调查、修复、重新开启和停止分开;一个晋升谓词要求每项检查既通过,又有证据支持其用于所请求的决策。实际教训是不仅要保留实验结果,还要保留它们对后续行动的限制。

英文摘要

Autonomous coding agents can remember an experiment yet carry forward a conclusion it does not justify. We reconstruct how evidence is reused in a 400-task NeuroGolf campaign, with selected wellbore-prediction records from the same operator as cross-domain comparisons. A numerical counterexample exposes an overbroad exclusion; other episodes distinguish what failed, incomplete and revised programs justify doing next. The wellbore records reveal an additional weakness: a coordinator correctly acknowledges a novelty-only rejection, then summarizes it as measured closure. Separately, a prefix-based acceptance gate at three thresholds yields worse target scores than ungated adaptation in the retained experiment. These cases motivate an inspectable handoff linking the tested proposition, its scope, candidate and evaluator identity, check statuses, and reopening condition. A dispatch tree separates investigation, repair, reopening and stopping; a promotion predicate requires that each check both passes and has evidence supporting its use for the requested decision. The practical lesson is to preserve not just experimental results, but their limits on subsequent action.

Comments31 pages, including technical appendices. Submitted to the NeurIPS 2026 Workshop on AI for Verifiable Coding (VeriCodeGen)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑