arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面包屑式搜索智能体

Breadcrumbing Search Agents

Xuebin Li, Hanqing Zhao, Siyuan Liang, Kejiang Chen, Weiming Zhang, Dacheng Tao, Nenghai Yu

arXiv 2608.04565首次发表:更新:

发表机构

School of Cyber Science and Technology, University of Science and Technology of China; College of Computing and Data Science, Nanyang Technological University(中国科学技术大学网络空间安全学院; 南洋理工大学计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM搜索智能体的安全问题,提出ACH和TGSE策略,通过操纵搜索结果形成连贯证据链,大幅提升攻击成功率。

AI 中文摘要

基于大语言模型(LLM)的搜索智能体被广泛用于信息检索任务,但它们依赖外部工具返回结果的特性引入了关键安全风险:执行过程中检索的网页内容不可信,使智能体面临提示注入和目标劫持的威胁。现有搜索智能体安全研究主要聚焦于静态网页内容注入,但现代智能体会发出后续查询并交叉验证不同来源,因此单次注入的页面往往会被稀释或拒绝。我们证明,传递搜索和页面观测结果的通道是脆弱的安全边界:除了使智能体暴露于单次中毒页面外,中介式搜索界面可反复引导智能体收集证据并形成最终答案。在受约束的工具中介威胁模型下,若在智能体轨迹中对证据进行协调,每次查询仅添加一个受控结果即可大幅提升攻击成功率。我们在该场景下研究了策略驱动的长程攻击系统,并提出权威链劫持(Authority-Chain Hijack, ACH),这是一种经专家优化的策略,可将孤立的搜索结果和页面内容操纵转化为跨看似相互印证来源的连贯证据链。ACH在所有基线中达到最高整体攻击成功率(ASR),在完整SafeSearch测试拆分上达到55.9% / 83.3%的ASR / MaxN ASR。我们进一步提出轨迹引导策略演化(Trace-Guided Strategy Evolution, TGSE),该方法可从执行轨迹自动改进攻击者策略,用轨迹驱动优化替代手动重新设计;其最强单设置在保留评估中达到71.4% / 95.0%的结果。

英文摘要

LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to prompt injection and goal hijacking. Prior work on search-agent safety primarily focuses on static web-content injection, but modern agents issue follow-up queries and cross-check competing sources, so a single injected page is often diluted or rejected. We show that the channel delivering search and page observations is a fragile security boundary: beyond exposing the agent to a single poisoned page, a mediated search interface can repeatedly steer how the agent gathers evidence and forms its final answer. Under a constrained tool-intermediary threat model, appending only one controlled result per query can substantially increase attack success when the evidence is coordinated across the agent's trajectory. We study this setting with a strategy-driven long-horizon attack system and introduce Authority-Chain Hijack (ACH), an expert-refined strategy that turns isolated search-result and page-content manipulations into a coherent evidence chain across seemingly corroborating sources. ACH achieves the highest Overall ASR among all baselines, reaching 55.9% / 83.3% ASR / MaxN ASR on the full SafeSearch test split. We further introduce Trace-Guided Strategy Evolution (TGSE), which automatically improves attacker strategies from execution traces, replacing manual redesign with trace-driven refinement; its strongest single setting reaches 71.4% / 95.0% in held-out evaluation.

Comments39 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑