发表机构
MIT CSAIL; Harvard Medical School; Managent(麻省理工学院计算机科学与人工智能实验室; 哈佛医学院; Managent)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对并发AI编码智能体的协调问题,提出非对称仓库谱系建模与验证器引导协调方法,通过MERGEGYM基准测试验证了谱系分诊在冲突预测和调度中的有效性,并强调在重放成本低时应验证一切。
AI 中文摘要
并发AI编码智能体产生了一个协调问题,其中廉价信号可能优先处理工作,但只有可执行的检查器才能确立所声称的属性。我们通过一个属性范围验证契约和MERGEGYM(一个包含开放时间范围预测、重放条件冲突解决和调度三个轨道的基准测试)来研究这种分离。在一个分层的715对谱系集(167个文本冲突)上,79个冲突(47.3%)发生在作者修改的文件集不相交的情况下。这是三路合并预期的一个操作上重要的代理不匹配:作者PR差异是针对PR特定基础测量的,而检查器将两个头与它们的共同合并基础进行比较。一个独立的谱系并集规则达到保留集AUROC 0.877;一个12特征逻辑模型达到0.882 [0.845, 0.917],PR-AUC为0.661,而一个未调优的随机森林在相同的决策时特征上达到0.902 [0.875, 0.928],PR-AUC为0.701。在33.3%的保留重放预算下,逻辑和随机森林模型分别恢复了81.2%和82.9%的冲突。补丁重建在48/79个零重叠冲突中成功,且所有48个都变得干净;其他31个案例不确定,因此这一检查验证了预期的三路合并解释,而不是声称新的Git机制。在T1中,零样本LLM达到AUC 0.704,LLM加元数据融合达到0.740。在T3中,决策时门在冻结标签重放下以65.0%的完工时间膨胀中位去重叠了91.7%的标记范围冲突。由于本地git merge-tree重放在我们的日志中已经便宜(中位数0.02秒),我们不声称谱系分诊仅节省此检查器:当精确重放便宜时,验证一切。本文中的所有实证保证仍限于文本可合并性或明确陈述的冻结标签调度目标。
英文摘要
Concurrent AI coding agents create a coordination problem in which cheap signals may prioritize work, but only an executable checker can establish the property being claimed. We study this separation through a property-scoped verification contract and MERGEGYM, a three-track benchmark for open-time scope forecasting, replay- conditioned conflict resolution, and scheduling. On a stratified 715-pair lineage set (167 textual conflicts), 79 conflicts (47.3%) occur despite disjoint authored file sets. This is an operationally important proxy mismatch expected from three-way merge: authored PR diffs are measured against PR-specific bases, whereas the checker compares both heads to their common merge base. A standalone lineage-union rule reaches held-out AUROC 0.877; a 12-feature logistic model reaches 0.882 [0.845, 0.917] with PR-AUC 0.661, and an untuned random forest on the same decision-time features reaches 0.902 [0.875, 0.928] with PR-AUC 0.701. At a 33.3% held-out replay budget, the logistic and random-forest models recover 81.2% and 82.9% of conflicts, respectively. Patch reconstruction succeeds for 48/79 zero- overlap conflicts and all 48 become clean; the other 31 cases are inconclusive, so this check validates the expected three-way-merge explanation rather than claiming a new Git mechanism. In T1, a zero-shot LLM reaches AUC 0.704 and LLM-plus- metadata fusion 0.740. In T3, a decision-time gate de-overlaps a median 91.7% of labeled scope collisions at 65.0% makespan inflation under frozen-label replay. Because local git merge-tree replay is already cheap in our logs (median 0.02 s), we do not claim that lineage triage saves this checker alone: when exact replay is cheap, verify everything. All empirical guarantees in this paper remain limited to textual mergeability or the explicitly stated frozen-label scheduling target.