arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示有帮助,但它们是否起到教学作用?评估代码生成中的技能迁移

Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

Will Badr

arXiv 2609.01106首次发表:更新:

AI 中文总结

该研究以HumanEval+、MBPP+为基准,测试Qwen2.5-3B-Instruct等模型,发现代码生成的相关提示多引导模型生成已有方案,未实现通用技能迁移。

AI 中文摘要

当提示将一个失败的生成程序转为通过的程序时,它是提供了缺失的信息,还是仅仅引导模型走向它本就可以生成的解决方案?我们使用可执行评估在HumanEval+和MBPP+上检验这些假设。对于Qwen2.5-3B-Instruct,自适应相关提示从79个选定的失败案例中挽救了36个;不相关提示挽救了19个,而8个无提示样本通过普通采样解决了46个问题,且恢复了相关提示挽救的36个中的31个。Phi-3.5-mini呈现相同模式:相关提示从101个失败案例中挽救了42个,不相关提示挽救了17个,无提示采样解决了57个,包括42个相关提示挽救的中的36个。由于提示条件使用不同的尝试预算,这些比较未分离出纯粹的语义效应。对Qwen的机制测试识别出由相关和不相关提示共享的稳定激活方向,持续添加该方向产生14次挽救和18次退化,无可检测的净准确率提升;学习到的低秩干预有正但不精确的估计效应。完整文本规范解决24个上下文定义问题中的22个,而测试的虚拟-KV前缀解决5-11个。生成后的隐藏状态探针跨基准迁移,合并AUROC为0.806和0.780,但其相对于令牌置信度的top-1选择优势在统计上未确定。总体而言,相关提示可挽救失败,但大多数被挽救的解决方案已可通过普通采样获得,且此处测试的内部干预未建立任务通用能力迁移。

英文摘要

When a hint turns a failing generated program into a passing one, does it provide missing information or merely steer the model toward a solution it could already produce? We test these hypotheses on HumanEval+ and MBPP+ using executable evaluation. For Qwen2.5-3B-Instruct, adaptive relevant hints rescue 36 of 79 selected failures; an unrelated hint rescues 19, while eight unhinted samples solve 46 and recover 31 of the 36 relevant-hint rescues. Phi-3.5-mini shows the same pattern: relevant hints rescue 42 of 101 failures, an unrelated hint rescues 17, and unhinted sampling solves 57, including 36 of the 42 relevant-hint rescues. Because the hint conditions use different attempt budgets, these comparisons do not isolate a purely semantic effect. Mechanistic tests on Qwen identify a stable activation direction shared by relevant and unrelated hints. Persistently adding this direction yields 14 rescues and 18 regressions, with no detectable net accuracy gain; learned low-rank interventions have a positive but imprecise estimated effect. Full textual specifications solve 22 of 24 context-defined problems, versus 5-11 for tested virtual-KV prefixes. Post-generation hidden-state probes transfer across benchmarks, with pooled AUROC 0.806 and 0.780, but their top-one selection advantage over token confidence is statistically unresolved. Overall, relevant hints can rescue failures, but most rescued solutions are already reachable through ordinary sampling, and the internal interventions tested here do not establish task-general capability transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑