发表机构
Efficient Computation Inc(高效计算公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨扩散语言模型内部激活是否编码代码正确性,通过线性探针在六个模型上区分成败,但探针引导未稳定提升生成性能,并分析了统计与可复现性局限。
AI 中文摘要
扩散语言模型通过反复更新部分掩码序列来生成代码。我们探究其内部激活是否编码了代码正确性,以及该信息能否改善生成过程。在六个扩散模型上,线性探针能够区分成功与失败的尝试,最强的读取信号通常出现在早期层之后。使用小型语义突变的对照实验支持了与正确性相关的联系,而非仅表面风格。在与模型置信度的比较中,探针点估计未提供一致的优势。在残差流中添加探针导出的方向在测试的引导设置中未产生可靠的改进,而相反方向则降低了性能。我们将这些观察与关于统计显著性或无法引导的一般性主张区分开来。补充方法、存档结果和代码记录了测试的干预措施及其统计校准和可复现性的局限性。
英文摘要
Diffusion language models generate code by repeatedly updating a partially masked sequence. We ask whether their internal activations encode code correctness and whether that information can improve generation. Across six diffusion models, linear probes distinguish passing from failing attempts, with the strongest reads generally appearing beyond the early layers. Controls using small semantic mutations support a connection to correctness rather than surface style alone. In comparisons with model confidence, probe point estimates offer no consistent advantage. Adding a probe-derived direction to the residual stream does not yield a dependable improvement in the tested steering settings, while the opposite direction degrades performance. We distinguish these observations from claims about statistical significance or a general inability to steer. Supplementary methods, archived results, and code document the tested interventions and the limits of their statistical calibration and reproducibility.