arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在神谕问题下对LLM测试生成中反馈驱动演化的审计与分解

Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

arXiv 2608.19626首次发表:更新:

AI 中文总结

该研究针对LLM测试生成中反馈驱动演化的神谕问题,通过多任务实验发现单一神谕会膨胀演化增益,提出审计-安慰剂协议以分离验证器人工制品等,为相关评估提供改进方案。

AI 中文摘要

执行反馈常被视为改进大语言模型(LLM)生成测试的自验证信号。然而,当生成的输入在单个已接受的程序上执行,且其输出被用作基准时,无效或未充分说明的输入会产生虚假的故障检测和表面上的演化增益。我们使用142个开发任务、114个锁定的外部任务和138个保留任务,结合两个代码模型、三个随机种子以及跨故障拟合的真实提交,对反馈驱动测试生成中的这种失败模式进行了审计。在三个已接受实现达成一致的外部输入上,生成的输出仅在27.79%和50.12%的案例中与专家组结果匹配。单一参考神谕使演化的测量增益膨胀了9.46至14.85个百分点;经审计后,同等预算的独立重采样比基于变异的演化表现好6.01至18.83个百分点。我们进一步将真正的三轮反馈循环与密度匹配的安慰剂进行比较,外部真实-安慰剂差异为+0.13和-0.50点,而保留集差异为+1.99和+0.28点,未提供细粒度反馈益处的有力证据。两名软件工程博士生进行的盲态语义审计将94.41%的专家组不确认输入归类为无效,但3.60%归类为有效,表明专家组分歧具有信息性但并非语义证明。我们提出一种审计-安慰剂协议,在对自演化测试生成器的评估中,将验证器人工制品、交互支架和基于基准的反馈信用分离开来。

英文摘要

Execution feedback is often treated as a self-verifying signal for improving LLM-generated tests. However, when generated inputs are executed on a single accepted program and its outputs are used as ground truth, invalid or underspecified inputs can create spurious fault detections and apparent evolutionary gains. We audit this failure mode in feedback-driven test generation using 142 development tasks, 114 locked external tasks, and 138 held-out tasks, with two code models, three seeds, and fault-cross-fitted real submissions. On external inputs for which three accepted implementations agree, generated outputs match the panel on only 27.79% and 50.12% of cases. A single-reference oracle inflates the measured gain from evolution by 9.46-14.85 percentage points; after auditing, equal-budget independent resampling outperforms mutation-based evolution by 6.01-18.83 points. We further compare a genuine three-round feedback loop with a density-matched placebo. External Real-Placebo differences are +0.13 and -0.50 points, while held-out differences are +1.99 and +0.28 points and do not provide robust evidence of fine-grained feedback benefit. A blinded semantic audit by two software engineering doctoral students classifies 94.41% of panel-disconfirmed inputs as invalid but 3.60% as valid, showing that panel disagreement is informative but not semantic proof. We propose an audit-and-placebo protocol that separates verifier artifacts, interaction scaffolding, and grounded feedback credit in evaluations of self-evolving test generators.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑