ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
ICON:基于推理时校正的代理间接提示注入防御
机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡) ; Alibaba(阿里巴巴) ; School of Computer Science, Peking University, China(计算机科学学院,北京大学,中国)
AI总结 ICON通过潜在空间轨迹探测和修正器实现对代理间接提示注入攻击的防御,有效提升任务连续性和安全性,同时保持高检测率和任务效用。
Comments 11 pages,