arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02035cs.CR

基于语义匹配的大语言模型智能体技能选择的隐式操纵

Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching

Qikai Wang, Yongzhao Zhang, Zhiwei Chen, Yimiao Sun, Jiguo Yu, Xiaosong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出ISM方法,通过三阶段策略隐式操纵LLM智能体的技能选择,在多域多模型上大幅提升目标选择率,且隐蔽性远优于显式引导,对多种防御工具仍有效。

中文摘要 AI 辅助

技能选择是大语言模型(LLM)智能体工作流程中的关键阶段,用于确定由哪个已安装技能处理用户请求。现有针对该阶段的攻击主要依赖显式提示注入或指令级引导,这会暴露可识别的操纵信号。本研究发现了技能选择的一个新的隐式攻击面:即使用户提示和技能描述单独来看是良性的,仍可通过策略性塑造它们的语义关系来偏向攻击者选定的技能。基于此观察,我们提出了基于语义匹配的隐式技能选择操纵方法(ISM),该方法联合塑造目标技能元数据和可复用提示,以在无显式选择指令的情况下操纵技能选择。具体而言,我们开发了一个三阶段策略,用于扩大语义覆盖范围、增强目标独特性并保留自然的提示措辞。在四个任务域和八个选择器模型上,ISM 将平均目标选择率(TSR)从 15.2% 提升至 63.5%;在配对比较中,ISM 达到 73.5% 的 TSR,仅比显式引导低 9.8 个百分点。人类评审员仅在 2.9% 的判断中阻止 ISM,而显式引导的这一比例为 91.4%;五个基于 LLM 的检查员对 ISM 的平均通过率为 82.9%,而显式引导仅为 37.4%。此外,ISM 对 PPL-W、Llama Prompt Guard 2 和 PIGuard 仍保持有效。

英文摘要

Skill selection is a key stage in LLM-agent workflows, determining which installed skill should handle a user request. Existing attacks on this stage primarily rely on explicit prompt injection or instruction-level steering, which can expose recognizable manipulation signals. In this work, we identify a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill. Based on this observation, we present Implicit Skill-Selection Manipulation via Semantic Matching (ISM), which jointly shapes target-skill metadata and reusable prompts to manipulate skill selection without explicit selection instructions. Specifically, we develop a three-stage strategy to broaden semantic coverage, strengthen target distinctiveness, and preserve natural prompt wording. Across four task domains and eight selector models, ISM increases the average target-selection rate (TSR) from 15.2% to 63.5%. In a matched comparison, ISM achieves a 73.5% TSR, only 9.8 percentage points below Explicit Steering. Human reviewers block ISM in only 2.9% of judgments, versus 91.4% for Explicit Steering, while five LLM-based inspectors pass ISM at an average rate of 82.9%, versus 37.4% for Explicit Steering. Moreover, ISM remains effective against PPL-W, Llama Prompt Guard 2, and PIGuard.

发表机构

  • University of Electronic Science and Technology of China(电子科技大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑