arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27674cs.SE

大语言模型能否解决真实的Java合并冲突?一项使用校准后的LLM作为评判者的评估

Can Large Language Models Resolve Real Java Merge Conflicts? An Evaluation with a Calibrated LLM-as-Judge

Bowen Shen

AI总结:

该研究针对真实Java合并冲突,构建仅用推理信号的LLM求解器,经校准的LLM评判者评估发现,其解决冲突的覆盖和匹配开发者方案的表现优于传统工具AutoMerge,但结构正确性仍需确定性检查保障。

AI中文摘要:

合并冲突是协同软件开发中反复出现的成本,传统的结构化和半结构化合并工具在处理这类冲突时经常弃权(不执行):当它们的启发式规则不适用时,会留下冲突未解决。相反,大语言模型(LLM)几乎可以为任何冲突生成候选解决方案,但大规模衡量这些解决方案是否真正良好却很困难,因为为每个模型输出获取人类期望的判断无法规模化。我们针对ConflictBench中的真实Java合并冲突,同时研究这两个问题。我们首先构建一个LLM求解器,作为仅使用推理时信号(冲突标记、Java解析器和重复声明检查)的生成-验证-重试智能体,且从未见过开发者的答案。然后我们用两套指标评估其解决方案:(1)作为G-Eval指标实现的开发者匹配LLM评判者,关键是在使用前已针对ConflictBench的人类标签进行校准;(2)不使用LLM的确定性结构有效性检查。在对292个人类标注案例的元评估中,该评判者达到100%的精度(零错误接受)和64.6%的召回率,因此每一次接受都是可信的,所有下游比率都是保守的下界。在这个经过验证的评判者下,LLM求解器在约55%的真实冲突上匹配开发者自己的解决方案(保守下限);在覆盖公平比较下,LLM(55-59%)比最强的传统工具AutoMerge(36.7%)高出约18-22个百分点;这种优势几乎完全来自覆盖范围,而非原始准确率,因为传统工具在20-90%的冲突上弃权,而强制解决下的LLM则不会弃权。最后,LLM评判者接受了5个未通过确定性结构检查的解决方案中的4个,这表明结构正确性不能委托给LLM。

英文摘要:

Merge conflicts are a recurring cost of collaborative software development, and the traditional structured and semi-structured merge tools that address them frequently abstain: when their heuristics do not apply, they leave the conflict unresolved. Large language models (LLMs) can instead produce a candidate resolution for almost any conflict, but measuring whether those resolutions are actually good at scale is hard, because obtaining human desirability judgments for every model output does not scale. We study both problems together on real Java merge conflicts from ConflictBench. We first build an LLM solver as a generate-validate-retry agent that uses only inference-time signals (conflict markers, a Java parser, and duplicate-declaration checks) and never sees the developer's answer. We then evaluate its resolutions with a two-metric suite: (1) a developer-match LLM-as-judge implemented as a G-Eval metric and, crucially, calibrated against ConflictBench's human labels before use, and (2) a deterministic structural-validity check that uses no LLM. On a meta-evaluation of 292 human-labeled cases, the judge reaches 100% precision (zero false accepts) at 64.6% recall, so every acceptance is trustworthy and every downstream rate is a conservative lower bound. Under this validated judge, LLM solvers match the developer's own resolution on about 55% of true conflicts (conservative floor), and under a coverage-fair comparison the LLMs (55-59%) beat the strongest traditional tool (AutoMerge, 36.7%) by roughly 18-22 points; the edge comes almost entirely from coverage, not raw accuracy, since the tools abstain on 20-90% of conflicts while the LLM under forced resolution abstains on none. Finally, the LLM judge accepted 4 of the 5 resolutions that fail the deterministic structural check, evidence that structural correctness must not be delegated to an LLM.

↑