发表机构
Systems and computer Engineering Carleton University Ottawa, Canada(加拿大渥太华卡尔顿大学系统与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出复杂度感知框架,利用圈复杂度等指标评估LLM代码理解,发现DeepSeek-Coder-V2和Llama在高复杂度代码上准确率显著下降,证明该评估比聚合准确率更具诊断性。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于需要理解现有源代码的软件工程任务,包括行为预测、函数解释、调试和代码审查。然而,聚合基准准确率可能掩盖模型可靠性随源代码结构复杂度增加而变化的情况。本文提出了一个复杂度感知框架,利用圈复杂度、嵌套深度、分支因子和 Halstead 体积来评估 LLM 的代码理解能力。我们通过两个互补任务评估 DeepSeek-Coder-V2 和 Llama:对 300 个 Python 函数进行自动输入输出预测,以及对 60 个函数的平衡子集进行人工评估的语义理解。这些函数被分为低、中、高三个复杂度区间。DeepSeek-Coder-V2 的总体自动准确率达到 78.33%,而 Llama 为 70.33%。然而,从低复杂度到高复杂度,准确率显著下降,DeepSeek-Coder-V2 从 93.52% 降至 52.78%,Llama 从 87.04% 降至 47.22%。错误预测始终与所有四个复杂度指标的较高值相关,相关性和逻辑回归分析证实了结构复杂度与正确性之间存在大致相当的负相关关系。人工语义理解显示出相同的退化模式,DeepSeek-Coder-V2 的准确率从 100.00% 降至 75.00%,Llama 从 90.00% 降至 60.00%。这些发现表明,复杂度感知评估比单独使用聚合准确率更能对 LLM 代码理解可靠性进行诊断性评估。
英文摘要
Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.