arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03267cs.SE

拒绝不可能:大语言模型中代码幻觉的分类与基准

Refusing the Impossible: A Taxonomy and Benchmark for Code Hallucination in Large Language Models

Vishnu Asutosh Dasu, Ashish Kundu, Gang Tan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究定义代码幻觉为无根据生成,提出三维分类体系,构建含270个提示的对抗性基准,发现多数模型在不可满足代码任务上易生成无根据代码,弃权率仅27%且无错误弃权。

中文摘要 AI 辅助

大语言模型(LLMs)常生成看似合理但不符合现实的代码,这些代码可能导入不存在的包,或声称实现违反已证明定理的算法,却仍能编译运行。本研究将代码幻觉定义为「无根据生成」,并将其与普通「代码错误」(有根据程序中的缺陷)区分开。我们提出包含三个维度的分类体系:「根据性」(绝对违反普遍真理与相对编造偶然或生态系统特定事实)、「表现层级」(句法、语义或事实)、「行为」(从自信编造到退化输出),并按严重程度排序。我们构建了一个包含故意不可满足任务的「对抗性」套件,其中正确响应是弃权(不执行),并按我们的分类体系对响应进行分类。该套件包含六种语言的270个提示、24个子类别,搭配91个匹配的可解决对照项,响应由经人工标签验证的两层协议评判(一致性达82%,κ=0.73)。在12个开放权重的代码与推理模型(4332个评判响应)中,模型在约60%的不可满足提示上生成无根据代码,仅在27%的提示上弃权(不执行),且在可解决对照项上未出现错误弃权(不执行)的情况。

英文摘要

Large language models (LLMs) often produce code that looks plausible but is not grounded in reality. The code may import packages that do not exist or claim to implement algorithms that violate proven theorems, while still compiling and running. We study \emph{code hallucination} as \emph{ungrounded generation} and separate it from ordinary \emph{code error} (bugs in otherwise grounded programs). We propose a taxonomy with three dimensions: \textbf{groundedness} (absolute violations of universal truths vs.\ relative fabrications of contingent or ecosystem-specific facts), \textbf{manifestation level} (syntactic, semantic, or factual), and \textbf{behavior} (from confident fabrication to degenerate output), organized into a severity ordering. We build an \textbf{adversarial} suite of deliberately unsatisfiable tasks where the correct response is to refuse and categorize the responses under our taxonomy. The suite contains \textbf{270 prompts} across six languages and 24 subcategories, paired with \textbf{91 matched solvable controls}, and responses are judged by a two-tier protocol validated against human labels (82\% agreement, $κ{=}0.73$). Across twelve open-weight code and reasoning models (4{,}332 judged responses), models produce ungrounded code on about 60\% of unsatisfiable prompts and refuse only 27\%, while wrongly refusing 0\% of the solvable controls.

发表机构

  • Pennsylvania State University(宾夕法尼亚州立大学)
  • Cisco Research(思科研究)

机构由 AI 辅助整理,请以论文原文为准。

↑