arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工具增强型大语言模型系统中的持久语义实体

Persistent Semantic Entities in Tool-Augmented LLM Systems

Zhaohui Wang

arXiv 2608.07952首次发表:更新:

AI 中文总结

该研究定义工具增强型LLM系统的持久语义实体(PSEs),证实其在24款不同规模模型中普遍存在,偏好/指令污染难自校正,上下文隔离自验证可降低污染,为智能体系统安全提供关键发现。

AI 中文摘要

工具增强型大语言模型智能体可在会话间保留隐式状态,通过事件激活并在智能体边界间传播,这一状态对标准调试工具基本不可见。我们将其形式化为持久语义实体(PSEs):由名称绑定、事件触发和跨边界传播定义的结构,并在来自11个家族的24个模型(参数规模15亿至1万亿)上对其进行评估。首先,所有测试模型均存在易感性(在20模型易感性面板上为20%至100%),其中名称绑定是必要且主导机制:若无名称绑定,污染度为0%。其次,持久性取决于污染类型而非模型规模或部署方式:偏好污染在所有测试模型上均无衰减地持续存在(t=10时为100%),指令污染在所有采用该机制的模型上持续存在,人格风格注入部分衰减(90%→10%),而事实注入具有模型依赖性——在Llama-3.1-8B和GPT-4o-mini上自校正,但在两个Qwen2.5-coder变体上保持峰值,因此我们不主张其普遍自校正。偏好和指令结果在我们的受控环境中跨供应商一致。第三,上下文隔离的自我验证在无专家参考的情况下实现20%至79%的减少(中位数36.5%),而基于关键词的检测会产生系统性误报,且污染沿四阶段智能体管道放大1.9倍(40%→75%)。偏好和指令污染具有持久性、缺乏自校正且难以被标准监控捕捉,是已部署智能体系统的特别值得关注的攻击面。

英文摘要

Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSEs): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B--1T parameters). First, every tested model is susceptible (20--100% on the 20-model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%$\to$10%), while factual injection is model-dependent---self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it self-corrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20--79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9$\times$ along a four-stage agent pipeline (40%$\to$75%). Preference and instruction contamination---persistent, lacking self-correction, and poorly captured by standard monitoring---represent a particularly concerning attack surface for deployed agent systems.

CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026). Camera-ready version; 32 pages . The openreview url is https://openreview.net/forum?id=zPjmtawzT8&noteId=CXB95ws5so and ICML poster url is https://icml.cc/virtual/2026/poster/60521

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑