arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型内部代码正确性表示的鲁棒性研究

On the Robustness of LLMs' Internal Representation of Code Correctness

Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez, Mahmoud Kassem, Sarah Nadi

arXiv 2608.08266首次发表:更新:

发表机构

New York University Abu Dhabi(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究系统探究大型语言模型内部代码正确性信号的鲁棒性,发现无最优提取配置且分离故障无法提升信号质量。

AI 中文摘要

现代语言模型生成的代码读起来很自然,但往往无法实现需求。研究表明,模型自身的置信信号与实际正确性校准不佳,这并不奇怪。评估正确性的一种有前景的方法是深入模型内部:通过对比正确和错误程序的隐藏状态,近期研究捕捉到一种代码正确性的内部信号,该信号无需执行测试,就能比模型的词级或陈述式置信度更好地判断候选解决方案。然而,该信号是通过一种特定方式捕捉的,留下一个重要问题:它是反映模型的鲁棒属性,还是该选择的人为产物。我们系统地研究这个问题,改变从模型内部提取信号的方式。此外,我们还探究信号质量是否受提取所用数据的限制,为此构建仅在导致错误的故障上存在差异的程序对。我们的结果表明,没有单一配置是最优的,且分离故障并无帮助。

英文摘要

Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness. A promising way to assess correctness looks inside the model: by contrasting the hidden states of correct and incorrect programs, recent work captured an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, with no test execution. However, this signal was captured under one particular way, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice. We study this question systematically, varying how the signal is extracted from the model internals. Besides this, we also ask if the signal's quality is limited by the data used to extract it, by constructing program pairs that differ only in the fault that makes them incorrect. Our results show that no single configuration is best, and that isolating the fault does not help.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑