arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代码语言模型是否利用测试?测试驱动代码生成的行为与表征研究

Do Code Language Models Follow Tests? Paired Interventions on Program Behavior

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

arXiv 2607.26244首次发表:更新:

AI 中文总结

该研究以Qwen2.5-Coder-7B和Qwen3.6-27B为对象,针对多个代码生成任务探究代码语言模型对测试的利用情况,发现测试通过语义引导和上下文扰动影响模型,仅表征变化不代表有效利用测试。

AI 中文摘要

公开测试被广泛用于引导大语言模型的代码生成,但目前尚不清楚模型将其视为可执行规范还是仅作为额外的提示上下文。本研究针对HumanEval+、MBPP+及近期LiveCodeBench任务,采用Qwen2.5-Coder-7B和Qwen3.6-27B模型开展测试驱动代码生成研究,对比仅自然语言提示、含相关可见测试、打乱输出、无关测试、仅断言测试、更强模型生成的合成测试等不同设置。评估结合隐藏/私有测试通过率、任务级行为翻转、线性探测及分层隐藏状态变化。可见测试可显著提升Qwen2.5在MBPP+上的性能,但对HumanEval+和LiveCodeBench的影响较小或不稳定;对于LiveCodeBench上的Qwen3.6,仅自然语言的通过率从13.1%升至39.4%,而相关原始测试仅提升2.9个百分点(p=0.458),合成high5测试较匹配基线仅提升1.7个百分点(p=0.701)。更多测试会引发更大的表征变化,但准确率提升会迅速饱和或消失;仅断言测试和无关测试提示可大幅改变隐藏状态却不提升正确性。控制提示长度后,任务级变化与有益及有害翻转均相关,且与有害翻转的关联显著更强。研究结果表明,测试通过语义引导和提示上下文扰动两种方式影响代码模型,仅表征变化无法证明模型有效利用测试。

英文摘要

Visible tests specify concrete program behavior, but an improvement in benchmark accuracy does not establish that a model follows the rule expressed by those tests. We study test utilization through matched prompting controls, paired semantic interventions, and test suites selected by fault detection. Our semantic intervention holds an underspecified description and its example inputs fixed while changing the correct outputs to express one of two valid rules. Evaluation on unseen inputs measures whether both generated programs follow their respective rules. Across five models and three runs of 120 paired instances from 20 specification families, mean switching rates range from 11.1\% to 65.8\%. Qwen3.8-27B has the highest point estimate, followed by Qwen3.6-27B at 60.6\%; their paired difference remains uncertain. Explicit descriptions elicit both rules from these two models on every instance, exposing a gap between implementation capability and adoption of test-specified rules. Correct expected outputs improve MBPP+ accuracy beyond inputs alone for all five models. On 180 tasks with fixed three-test suites, high-detection suites detect 30.3 percentage points more errors in a held-out pool dominated by reference mutants. The corresponding generation differences range from $-0.6$ to $+1.1$ points; all intervals include zero and remain compatible with some benefit. Paired interventions make test-specified rule changes measurable alongside implementation capability and benchmark correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑