AI 中文总结
该研究定义工具增强型LLM系统的持久语义实体(PSEs),证实其在24款不同规模模型中普遍存在,偏好/指令污染难自校正,上下文隔离自验证可降低污染,为智能体系统安全提供关键发现。
AI 中文摘要
工具增强型大语言模型智能体可在会话间保留隐式状态,通过事件激活并在智能体边界间传播,这一状态对标准调试工具基本不可见。我们将其形式化为持久语义实体(PSEs):由名称绑定、事件触发和跨边界传播定义的结构,并在来自11个家族的24个模型(参数规模15亿至1万亿)上对其进行评估。首先,所有测试模型均存在易感性(在20模型易感性面板上为20%至100%),其中名称绑定是必要且主导机制:若无名称绑定,污染度为0%。其次,持久性取决于污染类型而非模型规模或部署方式:偏好污染在所有测试模型上均无衰减地持续存在(t=10时为100%),指令污染在所有采用该机制的模型上持续存在,人格风格注入部分衰减(90%→10%),而事实注入具有模型依赖性——在Llama-3.1-8B和GPT-4o-mini上自校正,但在两个Qwen2.5-coder变体上保持峰值,因此我们不主张其普遍自校正。偏好和指令结果在我们的受控环境中跨供应商一致。第三,上下文隔离的自我验证在无专家参考的情况下实现20%至79%的减少(中位数36.5%),而基于关键词的检测会产生系统性误报,且污染沿四阶段智能体管道放大1.9倍(40%→75%)。偏好和指令污染具有持久性、缺乏自校正且难以被标准监控捕捉,是已部署智能体系统的特别值得关注的攻击面。
英文摘要
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSEs): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B--1T parameters). First, every tested model is susceptible (20--100% on the 20-model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%$\to$10%), while factual injection is model-dependent---self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it self-corrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20--79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9$\times$ along a four-stage agent pipeline (40%$\to$75%). Preference and instruction contamination---persistent, lacking self-correction, and poorly captured by standard monitoring---represent a particularly concerning attack surface for deployed agent systems.
CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026). Camera-ready version; 32 pages . The openreview url is https://openreview.net/forum?id=zPjmtawzT8¬eId=CXB95ws5so and ICML poster url is https://icml.cc/virtual/2026/poster/60521