发表机构
W. P. Carey School of Business, Arizona State University(亚利桑那州立大学W. P. 凯瑞商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型编写代码及测试用例的对抗性测试强化循环,通过机械预言机实现。实验发现早期跨谱系效应是工具工件,消除混淆因素后,同谱系评论家轮次有因果估计,跨供应商配置有差异,还发布了相关协议、记录和代码。
AI 中文摘要
大语言模型越来越多地编写代码及其测试用例,而代码覆盖记录的是运行的内容而非验证的内容。我们研究了在机械预言机下的对抗性测试强化循环:测试模型编写测试用例,变异测试标记注入缺陷后仍存活的代码,评论家模型编写测试用例来精确杀死这些代码,每个判定都由机械完成,不存在模型评判其他模型输出的情况。实验1中,在五个Python主题上(一个同谱系循环单元无法评分),该循环杀死了一次性生成遗漏的105个变异体且无一遗漏,跨谱系评论家问题返回了预先声明的无效结果。核心发现是一次剖析:早期分析报告的p = 9.5e - 66的跨谱系效应是工具工件,输出上限会悄悄截断冗长模型,仅通过对完整分析的对抗性审查才被发现。审查还发现了进一步的混淆因素,即每个分支都对自己的初始套件进行重采样;实验2消除了这一因素。在预注册的冻结共享轮0设计下(四个主题各进行五次重复,提前提交种子),同谱系评论家轮次杀死了冻结初始套件中剩余的78%的存活代码(平均增量杀死率0.783,95%聚类自举区间[0.592, 0.935]),这是一个重复内因果估计;跨供应商配置在手臂成本低5.5倍的情况下显示出正的试点差异(率差0.178,95%区间[0.039, 0.347];幅度由单个重复主导)。这比较了两种命名的模型 - 供应商 - 工具配置,而非孤立的谱系效应:差距的一部分是一种配置记录的操作失败,包括截断复发,现在已被检测和评分而非被掩盖。跨模型比较可能继承运行它们的工具的不对称性。我们发布了两个协议、所有记录和分析代码。
英文摘要
Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We study an adversarial test-hardening loop under a mechanical oracle: a Tester model writes tests, mutation testing names surviving injected defects, and a Critic model writes tests to kill exactly those, with every verdict decided mechanically, so no model judges another's output. In Experiment 1, on five Python subjects (one same-lineage-loop cell could not be scored), the loop killed 105 mutants that one-shot generation missed and lost none, and the cross-lineage-Critic question returned a pre-declared null. The central finding was an autopsy: an earlier analysis reported a cross-lineage effect at p = 9.5e-66 that was an instrument artifact, an output cap silently truncating the verbose model, caught only by adversarial review of the completed analysis. Review then found a further confound, each arm resampling its own initial suite; Experiment 2 removes it. Under a pre-registered frozen-shared-round-0 design (five replicates on each of four subjects, seeds committed in advance), same-lineage Critic rounds killed 78% of the survivors the frozen initial suite left standing (mean incremental kill rate 0.783, 95% cluster-bootstrap interval [0.592, 0.935]), a within-replicate causal estimate; the cross-provider configuration showed a positive pilot difference (rate gap 0.178, 95% interval [0.039, 0.347]; magnitude dominated by a single replicate) at 5.5x lower arm cost. This compares two named model-provider-harness configurations, not an isolated lineage effect: part of the gap is one configuration's receipted operational failures, including truncation recurrences, now detected and scored rather than laundered. Cross-model comparisons can inherit the asymmetries of the harness that runs them. We release both protocols, all receipts, and the analysis code.
Comments26 pages. Two pre-registered experiments; protocols, all run receipts, and analysis code at https://github.com/Jott2121/crucible