arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21656cs.SEcs.AI

跨模型大语言模型代码审查:应该用Claude审查Codex还是相反?

Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?

Zuodong Xiang, Yike Zhang, YueMing Zhang, Hailu Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究开发者同时使用Claude和Codex进行代码审查的成本、时间及配对顺序问题,通过对116个任务的六种条件实验发现,Claude审查Codex草稿效果好,反向则不佳,有用的配对是不对称的,应Claude审查Codex。

中文摘要 AI 辅助

开发者越来越多地同时使用两个编码代理,一个编写草稿,另一个进行审查。但不清楚这种配对是否值得投入成本和时间,以及配对顺序是否重要。我们对116个近期的中硬lc b任务进行了控制实验,涉及Claude和Codex在六种条件下,近似软件从业者的工作流程。结果表明,Claude审查Codex草稿能将通过率从71.6%提高到89.7%,Codex自我审查能提高到84.5%,而反向审查效果不佳。

英文摘要

Developers increasingly use two coding agents together: one writes a draft, and the other reviews it. However, it is not clear whether the pairing is worth its cost and time, or whether the order of the pairing matters. We run a controlled experiment on 116 recent hard and medium lcb tasks with Claude and Codex across six conditions to approximate a software practitioner's workflow: both solo baselines, both cross-model orderings, and both same-model orderings. The reviewer sees the problem and the writer's draft but cannot execute tests, which approximates a code review step. Claude review raises Codex drafts from 71.6% to 89.7% ($p_{BH}=.001$); Codex self review raises them to 84.5% ($p_{BH}=.022$). The reverse direction does not pay off: Codex reviewing Claude drafts drops the pass rate from 91.4% to 82.8% ($p_{BH}=.046$), and Claude self review leaves the 91.4% baseline unchanged. Our evaluation indicates that the useful pairing is asymmetric: use Claude to review Codex, not the other way around.

发表机构

  • University of California, Davis(加州大学戴维斯分校)
  • Johns Hopkins University(约翰霍普金斯大学)
  • California State University, Long Beach(长滩加州州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑