Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
你的智能体大模型会秘密编码间接提示注入暴露的潜在信号
专题命中 测试时计算 :reasoning(abstract);CoT(abstract_cn);分类 cs.AI
AI总结 本文研究智能体大模型的间接提示注入暴露问题,通过探测、防御、解释三方面分析,提出AGRI防御方法,可大幅降低攻击成功率且保持任务效用
Comments v2. Compared to v1, we include the Kimi-K3 model's results and conduct a series of new experiments on understanding probe generalization