AI 中文总结
本研究提出ElasticBack,通过耦合触发-规则优化在LLM智能体技能中植入条件性单技能后门,实验显示其攻击性能优异且隐蔽性强,可推动技能供应链的防御研究。
AI 中文摘要
智能体技能是大语言模型(LLM)智能体按需加载的指令与资源集合,形成了新兴的供应链,其中单个被投毒的技能会持续危害所有安装它的智能体。然而,现有的技能攻击要么对所有请求触发,要么依赖微调后的权重或多个技能,尚未探索条件性且低成本的后门。本研究提出ElasticBack,一种有效的条件性单技能后门,它在技能文档中植入规则R,并在用户查询中植入看似良性的触发T,仅当两者同时出现时恶意有效载荷才会触发。ElasticBack通过“触发即开关”结构将两者绑定,通过语义锚定规则注入生成R,随后冻结R并针对它用受隐蔽性约束的遗传搜索优化T,从而在保持后门无权重、在良性输入上处于休眠状态的同时,优化有效性与隐蔽性。在三个目标行为(各含50个技能)和四个智能体LLM上开展的大量实验表明,ElasticBack实现了高攻击成功率、近零误报率,且保留了干净准确率,可跨模型迁移,还能规避部署时的防御。这些结果推动了对技能供应链更强防御的研究。
英文摘要
Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill attacks either fire on every request or rely on fine-tuned weights or multiple skills, leaving a conditional and low-cost backdoor unexplored. In this work, we present ElasticBack, an effective conditional single-skill backdoor that plants a rule R in the skill document and a benign-looking trigger T in the user query, so the malicious payload fires only when both co-occur. ElasticBack binds the two sides through a trigger-as-switch construction, generating R via semantic-anchored rule injection. It then freezes R and evolves T against it with a stealth-constrained genetic search, so that effectiveness and stealth are optimized, keeping the backdoor weight-free and dormant on benign inputs. Extensive experiments across three target behaviors (50 skills each) and four agent LLMs show that ElasticBack attains a high attack success rate at a near-zero false-positive rate with preserved clean accuracy, transfers across models, and evades deployment-time defenses. These results motivate stronger defenses for the skill supply chain.