基于语义匹配的大语言模型智能体技能选择的隐式操纵
Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching
浏览论文内容
中文总结 AI 辅助
本研究提出ISM方法,通过三阶段策略隐式操纵LLM智能体的技能选择,在多域多模型上大幅提升目标选择率,且隐蔽性远优于显式引导,对多种防御工具仍有效。
中文摘要 AI 辅助
技能选择是大语言模型(LLM)智能体工作流程中的关键阶段,用于确定由哪个已安装技能处理用户请求。现有针对该阶段的攻击主要依赖显式提示注入或指令级引导,这会暴露可识别的操纵信号。本研究发现了技能选择的一个新的隐式攻击面:即使用户提示和技能描述单独来看是良性的,仍可通过策略性塑造它们的语义关系来偏向攻击者选定的技能。基于此观察,我们提出了基于语义匹配的隐式技能选择操纵方法(ISM),该方法联合塑造目标技能元数据和可复用提示,以在无显式选择指令的情况下操纵技能选择。具体而言,我们开发了一个三阶段策略,用于扩大语义覆盖范围、增强目标独特性并保留自然的提示措辞。在四个任务域和八个选择器模型上,ISM 将平均目标选择率(TSR)从 15.2% 提升至 63.5%;在配对比较中,ISM 达到 73.5% 的 TSR,仅比显式引导低 9.8 个百分点。人类评审员仅在 2.9% 的判断中阻止 ISM,而显式引导的这一比例为 91.4%;五个基于 LLM 的检查员对 ISM 的平均通过率为 82.9%,而显式引导仅为 37.4%。此外,ISM 对 PPL-W、Llama Prompt Guard 2 和 PIGuard 仍保持有效。
英文摘要
Skill selection is a key stage in LLM-agent workflows, determining which installed skill should handle a user request. Existing attacks on this stage primarily rely on explicit prompt injection or instruction-level steering, which can expose recognizable manipulation signals. In this work, we identify a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill. Based on this observation, we present Implicit Skill-Selection Manipulation via Semantic Matching (ISM), which jointly shapes target-skill metadata and reusable prompts to manipulate skill selection without explicit selection instructions. Specifically, we develop a three-stage strategy to broaden semantic coverage, strengthen target distinctiveness, and preserve natural prompt wording. Across four task domains and eight selector models, ISM increases the average target-selection rate (TSR) from 15.2% to 63.5%. In a matched comparison, ISM achieves a 73.5% TSR, only 9.8 percentage points below Explicit Steering. Human reviewers block ISM in only 2.9% of judgments, versus 91.4% for Explicit Steering, while five LLM-based inspectors pass ISM at an average rate of 82.9%, versus 37.4% for Explicit Steering. Moreover, ISM remains effective against PPL-W, Llama Prompt Guard 2, and PIGuard.
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。