发表机构
ILLC, University of Amsterdam(逻辑、语言与计算研究所,阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究代码模型内部编码类型信息的情况,利用Java和Python代码示例的并行数据集探测预训练代码模型的残差流获取类型表示,发现跨语言类型表示等结果,且结构对一些变化有鲁棒性,还发布了代码和数据集。
AI 中文摘要
当前最先进的代码模型虽性能出色,但内部编码类型信息的程度却鲜为人知。我们利用Java和Python代码示例的并行数据集,探测预训练代码模型的残差流以获取内部类型表示。结果表明,即使是无类型代码也会出现跨语言类型表示。此外,通过在一种语言上训练探测器来推断另一种语言中的参数和结果类型,测试隐藏状态是否线性编码类型化函数应用所隐含的结果类型。最后发现,这种结构对词汇扰动和跨语言句法变化具有一定鲁棒性。据我们所知,先前关于代码模型可解释性的工作尚未直接针对形式类型语义或跨语言类型表示。我们发布了代码和数据集。
英文摘要
State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly understood. We probe the residual streams of pretrained code models for internal type representations using a parallel dataset of Java and Python code examples. Our results show that cross-lingual type representations emerge even from untyped code. Moreover, we test whether hidden states linearly encode the result type implied by typed function application by training probes on one language to infer argument and result types in the other. Finally, we find that this structure is partly robust to lexical perturbations and cross-language syntactic variations. To the best of our knowledge, prior work on interpretability of code models has not directly targeted formal type semantics or cross-lingual type representations. We release our code and datasets.
Comments18 pages, 12 figures. Accepted at ESSLLI 2026 (StuS; double-blind)