arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估针对大语言模型生成代码中包幻觉的推理时防御措施

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

Alberick Euraste Djire, Iyiola E. Olatunji, Melissa Tessa, Earl T. Barr, Jacques Klein, Tegawendé F. Bissyandé

arXiv 2608.22652首次发表:更新:

发表机构

University of Luxembourg; AI4D (CITADEL); University College London(卢森堡大学; AI4D(CITADEL); 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对LLM生成代码的包幻觉问题,修正了评估方法,评估了七种推理时防御措施,提出包效用指标,发现对抗提示下RAG与Self-Refine表现更优,明确防御需匹配威胁模型与效用。

AI 中文摘要

大语言模型(LLM)越来越多地被用于代码生成,但它们经常会幻觉出不存在的软件包,这为软件供应链创造了可被利用的入口点。针对该问题,我们做出四项贡献:第一,我们表明,先前的评估方法会系统性地高估幻觉率,因为在某些语言中会将标准库模块错误归类为幻觉,对于Python而言,这种高估幅度达到9.4个百分点。第二,我们评估了七种用于缓解包幻觉的推理时防御措施,包括五种引导解码策略(Greedy、Contrastive、DoLa、Nudging和Active Layer-Contrastive Decoding)、一种迭代自精化方法(Self-Refine)以及一种基于检索增强生成(RAG)的防御措施。在涵盖五个模型家族和四种编程语言(Python、JavaScript、Ruby、Rust)的八个模型上,RAG在32种模型-语言配置中的18种配置下降低了包幻觉率(PHR)。第三,我们引入了包效用(Package Utility,PU)指标,用于评估防御措施是否保留了有效且与任务相关的建议。在被评估的策略中,Greedy解码提供了最强的平均缓解-效用权衡。第四,我们用植入了虚构包名的对抗性提示对所有策略进行压力测试,发现与标准提示相比,PHR飙升了多达45个百分点,其中Ruby始终是最易受攻击的语言(80.9%至95.2%)。在对抗性条件下,RAG和Self-Refine的表现优于所有仅解码的策略,这表明当提示具有主动攻击性时,稳健的防御需要外部依据或迭代自我验证。我们的研究结果将包幻觉重新定义为测量问题和解码时控制问题,并且表明防御措施的选择必须与威胁模型和建议效用相匹配。

英文摘要

LLMs are increasingly used for code generation, yet they frequently hallucinate non-existent software packages, creating exploitable entry points into the software supply chain. We make four contributions to this problem. First, we show that prior evaluation methodologies systematically inflate hallucination rates by misclassifying standard-library modules as hallucinations in some languages. For Python, the overestimation reaches 9.4 percentage points. Second, we evaluate seven inference-time defenses for mitigating package hallucinations, including five guided decoding strategies (Greedy, Contrastive, DoLa, Nudging, and Active Layer-Contrastive Decoding), an iterative self-refinement approach (Self-Refine), and a Retrieval-Augmented Generation (RAG)-based defense.. Across eight models spanning five families and four programming languages (Python, JavaScript, Ruby, Rust), RAG reduces the package hallucination rate (PHR) in 18 of 32 model--language configurations. Third, we introduce Package Utility (PU) to assess whether defenses preserve valid and task-relevant recommendations. Among strategies evaluated, Greedy decoding provides the strongest average mitigation--utility trade-off. Fourth, we stress-test all strategies under adversarial prompts seeded with fabricated package names and find that PHR surges by up to 45 percentage points relative to standard prompts, with Ruby consistently the most vulnerable language (80.9--95.2\%). Under adversarial conditions, RAG and Self-Refine outperform all decoding-only strategies, indicating that robust defense requires either external grounding or iterative self-verification when prompts are actively hostile. Our results recast package hallucination as both a measurement problem and a decoding-time control problem, and they demonstrate that the choice of defense must be matched to the threat model and recommendation utility.

CommentsAccepted ASE 2026

DOI:10.1145/3832783.3837555

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑