发表机构
Chengdu Institute of Computer Applications, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Institute of Multidisciplinary Research for Advanced Materials (IMRAM), Tohoku University; Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(中国科学院成都计算机应用研究所; 中国科学院大学; 东北大学先进材料多学科研究所; 深圳先进技术研究院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过可执行奖励分解和套件级度量,在规范驱动的测试生成中训练模型同时优化输入杀死与输出正确性,识别预言机转换瓶颈,并验证保留输入效用反馈的益处与局限。
AI 中文摘要
从自然语言规范生成测试既需要能够暴露错误行为的输入,也需要正确的预期输出。这两个要求未必会同步提升:模型可以通过选择更简单的输入来提高测试正确性,或者发现其预期输出无法预测的有用输入。我们通过可执行的奖励分解和套件级预言机转换度量来研究这种交互。我们的生成器在一次响应中联合生成五个输入-输出测试。在训练期间,经过审计的参考程序提供正确性反馈,而固定的错误程序库提供两个效用信号:潜在输入杀死和检查生成输出后的有效杀死。一个加性的GRPO目标保留了两个信号,无需在推理时执行。在包含506个训练任务和142个评估任务的审计TC-Bench子集上,三个独立训练的Qwen3.5-9B运行在步骤75时将完整测试正确性从28.59%提升至42.54%,输入杀死从24.06%提升至25.27%,有效完整杀死从12.23%提升至14.15%。匹配的50步消融研究揭示了一个权衡:移除杀死奖励会带来更高的正确性和略高的完整杀死,但将输入杀死降至21.60%。在64个训练池任务上的固定输入源-预言机交叉归因了主要的无杀死到完整杀死差异,这归因于更难的输入选择,而非在相同输入上的更差输出预测。这些结果将预言机转换识别为一个可度量的瓶颈,并展示了在联合测试生成中保留输入效用反馈的益处和局限。
英文摘要
Generating tests from a natural-language specification requires both an input that exposes faulty behavior and a correct expected output. These requirements need not improve together: a model can increase test correctness by choosing easier inputs, or discover useful inputs whose expected outputs it cannot predict. We study this interaction through executable reward decomposition and suite-level oracle-conversion measurement. Our generator jointly emits five input--output tests in one response. During training, audited reference programs provide correctness feedback, while a fixed bank of faulty programs provides two utility signals: potential input kill and effective kill after checking the generated output. An additive GRPO objective preserves both signals without requiring execution at inference time. On an audited TC-Bench split with 506 training and 142 evaluation tasks, three independently trained Qwen3.5-9B runs at step 75 increase full-test correctness from 28.59\% to 42.54\%, input kill from 24.06\% to 25.27\%, and effective full kill from 12.23\% to 14.15\%. Matched 50-step ablations reveal a trade-off: removing kill rewards yields higher correctness and slightly higher full kill, but lowers input kill to 21.60\%. A fixed-input source--oracle crossover on 64 training-pool tasks attributes the principal NoKill-to-FullKill difference to harder input selection rather than worse output prediction on identical inputs. These results identify oracle conversion as a measurable bottleneck and show the benefits and limits of preserving input-utility feedback in joint test generation.