arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编辑器只有只读权限:扩散语言模型中的正确性信号

A Heckler in the Hidden State: Correctness Signals in Diffusion Language Models

Angad Miglani, Samrath Singh Chadha, Kevin Li, Manas Venkata Sai Ravulapalli

arXiv 2609.36783首次发表:更新:

发表机构

Efficient Computation Inc(高效计算公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨扩散语言模型内部激活是否编码代码正确性,通过线性探针在六个模型上区分成败,但探针引导未稳定提升生成性能,并分析了统计与可复现性局限。

AI 中文摘要

扩散语言模型通过反复更新部分掩码序列来生成代码。我们探究其内部激活是否编码了代码正确性,以及该信息能否改善生成过程。在六个扩散模型上,线性探针能够区分成功与失败的尝试,最强的读取信号通常出现在早期层之后。使用小型语义突变的对照实验支持了与正确性相关的联系,而非仅表面风格。在与模型置信度的比较中,探针点估计未提供一致的优势。在残差流中添加探针导出的方向在测试的引导设置中未产生可靠的改进,而相反方向则降低了性能。我们将这些观察与关于统计显著性或无法引导的一般性主张区分开来。补充方法、存档结果和代码记录了测试的干预措施及其统计校准和可复现性的局限性。

英文摘要

Diffusion language models generate code by repeatedly updating a partially masked sequence. We ask whether their internal activations encode code correctness and whether that information can improve generation. Across six diffusion models, linear probes distinguish passing from failing attempts, with the strongest reads generally appearing beyond the early layers. Controls using small semantic mutations support a connection to correctness rather than surface style alone. In comparisons with model confidence, probe point estimates offer no consistent advantage. Adding a probe-derived direction to the residual stream does not yield a dependable improvement in the tested steering settings, while the opposite direction degrades performance. We distinguish these observations from claims about statistical significance or a general inability to steer. Supplementary methods, archived results, and code document the tested interventions and the limits of their statistical calibration and reproducibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑