arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

名称会造成危害:检测本地编码大语言模型中包名幻觉引发的slopsquatting风险

Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

Akash Raj, Sargam Sahu

arXiv 2608.23897首次发表:更新:

AI 中文总结

本文针对本地编码大语言模型的包名幻觉引发的slopsquatting风险,提出双层检测器结合LangGraph状态机,在300个提示词上实现76%无幻觉代码,用户满意度达4.4分,21/24用户有采用意向。

AI 中文摘要

当代码生成语言模型伪造Python包名时,若攻击者已在PyPI上预注册该名称,可将这种幻觉转化为供应链攻击,此事件被称为'slopsquatting'。本文提出一种双层检测器来应对该问题:第一层执行确定性PyPI存在性检查;第二层是随机森林分类器,基于从包名及其PyPI元数据中提取的10个特征进行训练。一个导入名称协调器连接两层,解决如'import cv2'与'pip install opencv-python'这类不会被安全绕过的情况。该检测器被嵌入LangGraph状态机,在温度参数升高时重试,若多次失败则路由至更强的后备模型。在300个精心设计的提示词上,该流程在76%的运行中生成无幻觉代码;主模型在28.7%的运行中耗尽重试预算,模型内重试恢复其中约四分之一的情况,跨模型后备进一步恢复剩余部分的16.5%。研究观察到四个发现:第一,被标记的幻觉中有一半是已在PyPI上注册的包,作为知名项目的低质量仿冒品,由分类器而非确定性层捕获(例如pil、faiss、tabula、haystack);第二,幻觉率随提示对抗性增强几乎线性上升,从常规编码场景的0-10%升至slopsquat诱饵场景的40-73%;第三,较弱的主模型在未辅助时拒绝了10个直接诱饵中的6个,表明近期的指令微调提供了基线防御;第四,当主模型与后备模型属于同一模型家族时,约84%的主模型失败会在后备模型上重现,这促使采用跨家族配对。一项包含24名参与者的用户研究显示,平均满意度为4.4分(满分5分),24人中有21人表示有采用意向。

英文摘要

When a code generating language model fabricates a Python package name, an adversary who has pre-registered that name on PyPI can convert that hallucination into a supply chain compromise. This event has been termed as 'slopsquatting'. We propose a two layer detector to counter this issue. The first layer performs a deterministic PyPI existence check. The second is a Random Forest classifier trained on ten features derived from the package name and its PyPI metadata. An import name reconciler bridges the two, resolving cases such as 'import cv2' versus 'pip install opencv-python' without a security bypass. The detector is embedded in a LangGraph state machine that retries at escalating temperatures and, on repeated failure, routes to a stronger fallback model. Across 300 curated prompts, the pipeline produces hallucination free code on 76% of runs. The primary exhausts its retry budget on 28.7%; intra model retries recover roughly a quarter of those, and cross model fallback recovers a further 16.5% of the remainder. Four findings have been observed. First, half of the flagged hallucinations are packages already registered on PyPI, as low quality lookalikes of well known projects, caught by the classifier rather than the deterministic layer (e.g., pil, faiss, tabula, haystack). Second, hallucination rate scales almost linearly with prompt adversariality, from 0 to 10% on routine coding to 40 to 73% on slopsquat baits. Third, the weaker primary refused 6 of 10 direct baits unaided, suggesting recent instruction tuning provides a baseline defense. Fourth, when primary and fallback share a model family, approximately 84% of primary failures recur on the fallback, motivating cross family pairing. A user study (n = 24) reports mean satisfaction 4.4 out of 5 and 21 of 24 stated adoption intent.

Comments14 pages, 2 figures, 6 tables. Code and data at https://github.com/sargamsahu1011/package-hallucination-detector

DOI:10.5281/zenodo.22087562

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑