AI 中文总结
OpenCodeReview针对LLM代码评审智能体的非确定性与上下文局部性缺陷,通过在三个流水线节点注入确定性,在AACR-Bench上实现SEM-F1提升且大幅降低令牌消耗,性能优于主流编码智能体。
AI 中文摘要
基于大语言模型(LLM)的代码评审智能体有望实现可扩展、全天候的评审服务,但当前系统存在两个相互交织的缺陷:(1)非确定性——无约束的工具使用导致评审结果不稳定;(2)上下文局部性——评审者仅能访问代码差异(diff),限制了可发现问题的深度。这两个缺陷催生了三大挑战:上下文检索错位、多文件拉取请求(PR)中存在一致性-效率权衡、以及会削弱信任的幻觉评论。为解决这些问题,我们提出OpenCodeReview,其构建于针对不确定性智能体的确定性工程之上:我们不给予智能体最大自由度,而是在三个明确的流水线节点注入确定性。规则引导调度(Rule-Guided Dispatch)采用多层规则系统确定性地选择文件和评审标准,消除智能体驱动分类的变异性;基于基础的文件评审(Grounded File Review)将自由探索替换为通过ReAct循环暴露的精选工具集,同时文件级别的并行子智能体(SubAgents)平衡上下文一致性与效率,并按需恢复跨文件依赖;独立反思(Independent Reflection)在非对称信息边界下引入优先证伪的过滤器——反思器仅能看到代码差异,无法访问智能体的工具增强探索——从而消除幻觉评论且不会产生自增强偏差,在保留召回率的同时提升了精确率。在AACR-Bench(含200个真实PR、10种语言、1505条专家验证评论)上,OpenCodeReview在六个大语言模型后端上均优于主流编码智能体(如Claude Code和Codex),实现了最高2.17倍的SEM-F1(25.10%对比11.57%),同时消耗的令牌数减少了5-15倍。我们在该URL开源OpenCodeReview。
英文摘要
LLM-based code review agents promise scalable, always-on review, yet current systems suffer from two intertwined weaknesses: (1) non-determinism--unbounded tool use makes review outcomes unstable, and (2) context locality--the reviewer's access remains bounded to the diff, capping discoverable issue depth. Both give rise to three challenges: misaligned context retrieval, a coherence-efficiency trade-off in multi-file pull requests, and hallucinated comments that erode trust. To address these, we introduce OpenCodeReview, built on deterministic engineering for uncertain agents: rather than granting maximal freedom, we inject determinism at three deliberate pipeline points. Rule-Guided Dispatch uses a multi-layer rule system to deterministically select files and review criteria, eliminating variability of agent-driven triage. Grounded File Review replaces free-form exploration with a curated tool set exposed through a ReAct loop, while file-level parallel SubAgents balance context coherence against efficiency and recover cross-file dependencies on demand. Independent Reflection introduces a falsification-first filter under an asymmetric information boundary--the reflector sees only the diff, not the agent's tool-augmented exploration--removing hallucinated comments without self-reinforcing bias, improving precision while preserving recall. On AACR-Bench (200 real-world PRs, 10 languages, 1,505 expert-verified comments), OpenCodeReview outperforms mainstream coding agents (e.g., Claude Code and Codex) across six LLM backends, achieving up to 2.17x higher SEM-F1 (25.10% vs. 11.57%) while consuming 5-15x fewer tokens. We open-source OpenCodeReview at https://github.com/alibaba/open-code-review.