arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23961cs.SEcs.AIcs.CL

评估语言模型在跨语言代码功能等价性上的表现

Evaluating Language Models on Cross-Language Code Functional Equivalence

  • North Carolina State University(北卡罗来纳州立大学)
  • Federal University of Ceará(塞阿拉联邦大学)
  • Federal University of Campina Grande(大坎皮纳联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Hui Sun, Anderson Uchôa, Rohit Gheyi, Wesley K. G. Assunção

中文总结 AI 辅助

本研究构建PolyHuman跨语言代码数据集,评估GPT-o4-mini等LLMs的代码功能等价性判断能力,发现其存在难度依赖错误、语言敏感性及运行不稳定等局限,无法可靠捕获功能等价性。

中文摘要 AI 辅助

背景:大型语言模型(LLMs)在各类代码理解任务中展现出强大性能,使许多人认为它们能够对程序语义进行推理。然而,现有评估主要聚焦于单语言场景或依赖人工生成的代码,引发了当前结果是否反映真实语义理解的担忧。目标:本研究探究LLMs能否准确判断人类编写的不同编程语言代码间的功能等价性,该场景需要超越表面相似性的深层推理。方法:我们引入PolyHuman数据集,包含C++、Java和Python三种编程语言的人类编写程序。利用该数据集,我们评估开放权重和专有LLMs的语言内及跨语言等价性检测,选择GPT-o4-mini作为代表性模型评估其稳定性。随后,我们手动分析81个模型错误判断功能等价性的系统性分歧案例,检查代码逻辑和生成的思维链(Chain-of-Thought)推理。最后,我们对这些失败案例进行分类,并对比GPT-o4-mini、Claude-Opus-4.7和Gemini-3-Flash的情况,以确定它们是否反映模型特定问题或当前最先进LLMs的更广泛局限。结果:我们发现等价性判断存在难度依赖的 breakdown(更难的问题会使模型越来越倾向于将非等价代码错误分类为等价),表现最佳的模型对编程语言存在模型特定敏感性(尤其在Python上表现出更保守的行为),且部分依赖基于相似性的线索。GPT-o4-mini在相同设置下还表现出显著的运行间不稳定性,表明其能力不一致而非完全缺失。结论:当前LLMs无法可靠捕获语言内或跨语言的功能等价性。

英文摘要

Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to believe that they can reason about program semantics. However, existing evaluations primarily focus on single-language settings or rely on synthetically generated code, raising concerns about whether current results reflect true semantic understanding. Aims: We investigate whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity. Method: We introduce PolyHuman, a dataset of human-written programs in CPP, Java, and Python. Using this dataset, we evaluate intra- and inter-language equivalence detection across open-weight and proprietary LLMs, selecting GPT-o4-mini as a representative model to assess stability. We then manually analyze 81 cases of systematic disagreement in which models incorrectly judge functional equivalence, examining the code logic and the generated Chain-of-Thought reasoning. Finally, we categorize these failures and compare them across GPT-o4-mini, Claude-Opus-4.7, and Gemini-3-Flash to determine whether they reflect model-specific issues or broader limitations of state-of-the-art LLMs. Results: We identify a difficulty-dependent breakdown in equivalence judgment (harder problems make the model increasingly prone to misclassifying non-equivalent code as equivalent), a model-specific sensitivity to programming language for the best-performing model (particularly a more conservative behavior on Python), and a partial reliance on similarity-based cues. GPT-o4-mini also shows substantial run-to-run instability under identical settings, indicating inconsistent rather than absent capability. Conclusions: Current LLMs do not reliably capture functional equivalence within or across languages.

补充信息

↑