arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.13962cs.SE

CodeGlance:通过多维特征分析理解LLM中的代码推理挑战

CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis

Yunkun Wang, Xuanhe Zhang, Junxiao Han, Chen Zhi, Shuiguang Deng

更新

AI总结:

CodeGlance通过多维特征分析揭示LLM在代码推理中的挑战,发现未见过的函数推理对小型模型尤为困难,提出改进策略并提供开发指导。

AI中文摘要:

在现代软件开发中,开发者经常需要快速理解代码行为——无论是审查拉取请求、调试问题还是导航不熟悉的代码库。这种对动态程序行为的推理能力是有效软件工程的基础,并越来越多地由大语言模型(LLMs)支持。然而,现有的代码推理研究主要集中在孤立的代码片段上,忽视了涉及外部API交互和不熟悉函数的现实场景的复杂性。这种差距阻碍了我们对LLM在不同编程背景下代码推理挑战本质的理解。我们提出了CodeGlance,一个多维基准,调查代码推理挑战在三个现实场景中的情况:内在逻辑推理、API交互推理和未见过的函数推理。通过系统评估7种最先进的LLM,我们发现未见过的函数推理对小型模型来说是一个重大挑战,Qwen2.5-3b在未见过的函数上仅达到6.0%的准确率,而在熟悉API上为37.5%。我们识别出关键的代码复杂性特征——包括执行轨迹长度、API调用次数和控制流复杂性——这些特征在不同场景中显著影响代码推理难度。我们进一步研究了常见的增强策略,包括CoT、文档检索和代码搜索,如何提高推理性能,发现它们的有效性在挑战源于逻辑复杂性还是知识缺口时差异显著。这些发现为开发更强大的代码推理系统和在现实软件开发中部署基于LLM的编程助手提供了可行的指导。

英文摘要:

In modern software development, developers frequently need to understand code behavior at a glance -- whether reviewing pull requests, debugging issues, or navigating unfamiliar codebases. This ability to reason about dynamic program behavior is fundamental to effective software engineering and increasingly supported by Large Language Models (LLMs). However, existing studies on code reasoning focus primarily on isolated code snippets, overlooking the complexity of real-world scenarios involving external API interactions and unfamiliar functions. This gap hinders our understanding of what truly makes code reasoning challenging for LLMs across diverse programming contexts. We present CodeGlance, a multi-dimensional benchmark investigating code reasoning challenges across three realistic scenarios: intrinsic logic reasoning, API interaction reasoning, and unseen function reasoning. Through systematic evaluation of 7 state-of-the-art LLMs, we reveal that unseen function reasoning poses significant challenges especially for smaller models, with Qwen2.5-3b achieving only 6.0\% accuracy on unseen functions compared to 37.5\% on familiar APIs. We identify critical code complexity features -- including execution trace length, API invocation count, and control flow complexity -- that significantly impact code reasoning difficulty across scenarios. We further investigate how common augmentation strategies, including CoT, document retrieval, and code search, can improve reasoning performance, finding that their effectiveness varies substantially depending on whether challenges stem from logical complexity or knowledge gaps. These findings provide actionable guidance for developing more capable code reasoning systems and deploying LLM-based programming assistants in real-world software development.

↑