发表机构
Daejeon Jungang Cheonggua Co., Ltd.(大田中央青果有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验探究跨模型评审在LLM验证中的效用,发现其与同模型评审各有优劣,但二者结合可提升错误匹配率。
AI 中文摘要
大语言模型现已能生成代码、文档和分析,并越来越多地被用于评审此类输出。我们探究了由不同模型进行第二次评审何时会有帮助。基于作者早先的预印本(其中在同一模型内变化了上下文、重复和角色结构),我们在一个受控实验中测试了模型独立性:30个工件,包含150个植入错误,10种评审条件,以及900次评审会话,涉及来自两家开发者的三个评审模型。在该实验中,(1)顶级跨模型评审者在F1分数上与同一模型在新会话中的评审(CCR)无显著差异,但这并未确立等价性;(2)两者发现的错误部分不同(Jaccard系数为41.2%);以及(3)在两次评审调用时,一次CCR加上一次跨模型评审匹配到的植入错误多于两次CCR评审(56.7%对42.7%;Holm校正后p=0.006),但并未显著多于由顶级跨模型评审者进行的两次评审,因此模型差异与评审者能力未能分离。一个轻量级跨模型评审者的得分不高于同模型评审。对评审者隐瞒需求会提高两个较低层级评审者的F1分数,但对最高层级无此效果,这些未经验证的估计值模式取决于失败会话的计分方式。在分析之前,我们审计了所有会话记录,排除了一个来源不明的基线运行和14次失败调用;包含所有会话的结果也已报告。对另一基准的公共检测器输出的部分检查既未复现也未反驳主要比较结果。记录、工件和脚本可向作者索取。
英文摘要
Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.
Comments16 pages, 2 figures, 7 tables. Follow-up to arXiv:2603.12123 and arXiv:2603.21454. v2: corrects two condition labels in Table 1 (CCR sees the artifact only; SA runs in a new session) and dependent interpretations; adds review prompts, a TP/FP breakdown by severity, and limitations; states how each reviewer was run; softens case studies. Numbers unchanged except removed B5 percentages