arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27990cs.CRcs.AI

CAITLYN:大语言模型智能体能否自主合成针对新型注入攻击的防御措施?

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体注入攻击防御的三难困境,提出智能体无关防御中间件CAITLYN,其含即时防御与自主合成新防御的双系统,在标准基准与新基准Emerging上均表现优异,可降低攻击成功率。

中文摘要 AI 辅助

针对大语言模型(LLM)智能体的提示注入攻击,旨在将恶意指令或内容引入智能体检索到的外部文本源,迫使底层LLM执行超出其良性范围的有害操作。尽管当前防御措施能有效应对已知注入攻击,但由于攻击变体和新兴威胁,将其部署到LLM智能体环境仍具挑战性。此外,现有解决方案通常存在内在的三难困境,即运行时效率、上下文精度和适应性之间的持续权衡。为解决这一差距,我们提出了连续注入威胁智能体防御中间件CAITLYN(Continuous Agents for Injection Threats via Lifelong Yielding Nexus),它是一种与智能体无关的防御中间件。CAITLYN集成了两个系统:系统I专注于使用两层库对现有攻击进行即时防御,第0层为基于规则的检测脚本,第1层为优化的基于LLM的精准推理;系统II则用于监测潜在异常信号并尝试合成新的防御措施。在标准基准测试中,CAITLYN的检测性能与最先进的防御措施相当,且比LLM-as-a-judge基线的令牌开销更低。在我们新推出的、包含新型注入技术的感知交付基准Emerging上,静态基线和单独的系统I配置仍易受攻击,而系统II可自主合成经验证的防御能力,在三种不同的智能体环境中大幅降低了攻击成功率。

英文摘要

Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known injection attacks, deploying them in LLM agent environments remains challenging due to attack variants and emerging threats. Moreover, existing solutions typically suffer from an inherent trilemma, i.e., a constant trade-off among runtime efficiency, contextual precision, and adaptability. To bridge this gap, we propose Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware. CAITLYN integrates two systems. System I focuses on immediate defense against existing attacks using a two-tiered library: Tier-0 for rule-based detection scripts and Tier-1 for optimized LLM-based accurate inference. System II, in contrast, is deployed to monitor potential abnormal signals and attempt to synthesize new defenses. On standard benchmarks, CAITLYN matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines. On Emerging, our new delivery-aware benchmark featuring novel injection techniques, static baselines and the standalone System I configuration remain vulnerable. In contrast, System II autonomously synthesizes verified defense capabilities, substantially lowering the attack success rate across three diverse agent environments.

补充信息

↑