arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用上下文增强的代码表示检测参数交换错误

Detecting Argument-Swap Bugs Using Context-Enhanced Code Representations

Subrata Das, Ali Aman, Muhammad Asaduzzaman, Kawser Wazed Nafi, Salimur Choudhury

arXiv 2609.17844首次发表:更新:

发表机构

Lakehead University; University of Windsor; Polytechnique de Montreal; Queen’s University(拉凯德湖大学; 温莎大学; 蒙特利尔高等理工学院; 女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BugProbe,一种无需调用到定义映射的基于学习的方法,利用上下文增强的代码表示检测Python方法调用中的参数顺序错误,在合成与真实数据集上取得高准确率并优于基线。

AI 中文摘要

源代码元素的名称传达了丰富的语义信息,并已广泛应用于软件工程任务,如错误检测、代码补全、类型预测和代码分类。先前的研究利用方法参数与形参名称之间的词汇相似性来检测由参数顺序错误引起的缺陷,通常依赖于建立方法调用与其对应定义之间的映射。然而,在像Python这样的动态类型语言中,这种映射往往难以获得。在本文中,我们提出了BugProbe,一种基于学习的方法,用于检测Python方法调用中参数顺序错误,该方法不需要调用到定义的映射。我们的方法利用多种上下文信息源,包括局部上下文和参数使用上下文,并将基于名称的相似性与机器学习相结合,以构建方法参数的表现性表示。我们从GitHub上星标数前1000的仓库中收集了一个包含132,739个Python源文件的新数据集,产生了3,371,244个合成训练示例,并贡献了一个从提交历史中人工验证的55个真实世界参数交换错误的精选基准。我们在该数据集上评估了我们的方法,并表明它在标准评估指标上实现了高准确性,并且始终优于最先进的基线方法。这些结果表明,在不依赖显式调用到定义解析的情况下,有效检测参数顺序错误是可能的,使得该方法非常适合动态类型语言环境。

英文摘要

Names of source code elements convey rich semantic information and have been widely used in software engineering tasks such as bug detection, code completion, type prediction, and code classification. Prior studies exploit lexical similarity between method arguments and formal parameter names to detect bugs caused by incorrectly ordered arguments, typically relying on establishing mappings between method calls and their corresponding definitions. However, such mappings are often difficult to obtain in dynamically typed languages like Python. In this paper, we present BugProbe, a learning-based approach for detecting incorrectly ordered arguments in Python method calls that does not require call-to-definition mappings. Our approach leverages multiple sources of contextual information, including local context and argument usage context, and combines name-based similarity with machine learning to construct expressive representations of method arguments. We collect a new dataset of 132,739 Python source files from the top-1,000 starred GitHub repositories, yielding 3,371,244 synthetic training examples, and contribute a curated benchmark of 55 real-world argument-swap bugs manually verified from commit histories. We evaluate our approach on this dataset and show that it achieves high accuracy and consistently outperforms a state-of-the-art baseline across standard evaluation metrics. These results demonstrate that effective detection of argument-ordering bugs is possible without relying on explicit call-to-definition resolution, making the approach well suited for dynamically typed language settings.

CommentsAccepted in the 26th IEEE International Conference on Source Code Analysis and Manipulation (SCAM 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑