不存在的技能:对大语言模型代理中幻觉技能推荐的大规模研究
Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
浏览论文内容
中文总结 AI 辅助
研究大语言模型代理中技能名称幻觉漏洞,通过大规模测量发现各配置均有此问题,系统生成众多幻觉名称且非随机,测试的模型级防御存在安全与可用性冲突,修复需全生态系统结构改变。
中文摘要 AI 辅助
大语言模型代理通过从开放注册表下载技能来获取新功能。开发者通常让代理推荐并安装技能,而非手动浏览目录。这带来风险:代理常为注册表中不存在的技能虚构名称,即技能名称幻觉。虚假名称看似无害,却为供应链攻击打开大门。对手可促使代理返回虚假名称,预注册恶意技能。我们首次对技能名称幻觉进行大规模测量,评估12种配置下的15000个提示。结果显示这是系统漏洞,各配置均有幻觉,独立大语言模型平均幻觉率为36.0%,代理为36.9%,实际开发者问题中升至43.1%。系统共生成5669个不同的幻觉名称,且非随机。代理在不同提示和模型中重复相同虚假名称,为攻击者提供可靠目标。我们测试了四种模型级防御,发现安全与可用性严重冲突。最强的检索基础防御将幻觉率从40.8%降至3.2%,但实用性严重受损。因此,技能名称幻觉是极易利用的漏洞,修复需全生态系统结构改变,如注册表级名称预留和经过验证的推荐管道。
英文摘要
LLM agents acquire new capabilities by downloading skills from open registries. Instead of browsing these catalogs manually, developers typically ask the agent to recommend and install a skill. This convenience hides a risk: agents frequently invent names for skills that exist in no registry. We term this flaw skill name hallucination. A fake name may seem harmless, but it opens the door to supply-chain attacks. Because registries rarely verify publishers, an adversary can prompt the agent, collect the fake names it returns, pre-register malicious skills under them, and wait for a victim to install the payload. We conducted the first large-scale measurement of skill name hallucination, evaluating 15,000 prompts across 12 configurations (4 standalone LLMs and 8 agents). We conservatively counted a name as hallucinated only if it was missing from all live registries and GitHub. The results reveal a systemic vulnerability: every configuration hallucinates. Rates average 36.0% for standalone LLMs and 36.9% for agents, rising to 43.1% on real-world developer questions. In total, the systems generated 5,669 distinct hallucinated names. Crucially, these names are not random noise. Agents repeat the same fake names across prompts and models, giving attackers highly reliable targets to hijack. Finally, we tested four model-level defenses and found a severe conflict between security and usability. The strongest, retrieval grounding, cut the hallucination rate from 40.8% to 3.2% but crippled usefulness: even the best-defended system recommended the correct skill only about one in six times. Skill name hallucination is thus a highly exploitable vulnerability requiring minimal attacker effort. Fixing it cannot rely on prompt engineering or model tuning alone. It demands ecosystem-wide structural changes: registry-level name reservations and verified recommendation pipelines.