发表机构
The University of Texas at Austin; Cisco Research(德克萨斯大学奥斯汀分校; 思科研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出程序可执行性预测(PrEx)任务,构建含系统生成非法变换的数据集,评估开源编码LLM,发现其依赖预训练先验而非给定语义,在修改语义及复杂程序上表现差。
AI 中文摘要
大语言模型(LLM)已在代码生成、代码翻译等各类软件工程任务中展现出出色能力,但其性能的一项关键局限可能在于对编程语言语义的理解不足。即便提供了明确的语义规则,目前仍不清楚LLM是应用了这些规则,还是依赖预训练期间学到的先验知识。我们以一项新任务——程序可执行性预测(PrEx),探究LLM是依赖先验知识还是给定语义,该任务要求模型在给定程序的语法和操作语义的前提下,预测程序在语义上是合法还是非法(若非法,需指出其违反了哪条形式规则)。由于PrEx任务需要合法和非法程序,我们基于合法程序构建了包含系统生成的非法变换的数据集。我们针对开源编码LLM,在两种语义形式体系、两种语义偏移下,对人工编写、LLM翻译、模糊测试器生成的程序划分进行评估。研究发现,LLM依赖预训练先验而非系统应用给定规则,在修改后的语义上表现尤其差,且随着程序复杂度提升性能进一步下降。PrEx任务可在指定网址获取。
英文摘要
Large language models (LLMs) have shown proficiency in various software engineering tasks, such as code generation and translation. However, a key limitation in their performance may be their (lack of) understanding of programming-language semantics. Even when explicit semantics are given, it remains unclear whether LLMs apply those rules or lean on priors learned during pre-training instead. We study if LLMs lean on priors or given semantics with a novel task--Program Executability Prediction (PrEx)--that asks models to predict whether a program is semantically valid or invalid (and, if invalid, which formal rule it violates) given the program's syntax and operational semantics. Because PrEx requires both valid and invalid programs, we build a dataset with systematically generated invalid transformations derived from valid programs. We evaluate open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits. Our findings show that LLMs lean on pre-training priors rather than systematically applying the given rules, performing especially poorly on modified semantics and degrading further as program complexity increases. PrEx is available at https://github.com/EngineeringSoftware/prex.
CommentsAccepted at LMPL 2026