AI 中文总结
MemSentry是一个配置驱动的框架,通过拦截持久性记忆写入并基于源信任、语义风险等做出接受、审查或隔离决策,以检测智能体AI中的记忆投毒攻击。
AI 中文摘要
具有持久性记忆的智能体AI系统引入了一种独特的攻击面,称为记忆投毒攻击,其中对抗性构造的内容被存储在长期记忆中,并随后影响智能体的未来行为。此类攻击可以抑制安全警报、促进权限提升、改变信任关系或覆盖安全策略,而无需修改底层模型权重或系统提示。为了应对这一威胁,我们提出了MemSentry,一个正式的、配置驱动的框架,它拦截提议的持久性记忆写入,并产生确定性的接受、审查或隔离决策。MemSentry通过联合考虑源信任、语义风险、组件依赖DAG上的攻击半径、访问风险以及一个带符号的安全状态增量(该增量捕获操作是削弱还是加强系统的安全态势)来评估每次写入。我们使用一个包含20个资产的随机依赖DAG和一个10×20的用户访问控制矩阵来实例化受保护环境,并使用分层70/30训练/测试划分,在1,000个由GPT-4生成的场景上评估该框架。语义分类被视为一个可插拔组件而非主要贡献,我们比较了四种代表性方法:基于规则的Regex、TF-IDF+SVM、SBERT+LR和SetFit。SBERT+LR取得了最佳整体性能,准确率为91.7%,宏F1分数为0.908,而所有四种方法都能检测到100%的外部隔离类威胁。对于经过验证的内部人员,当源信任最大(T=1)时,MemSentry不会自动隔离可疑操作,而是将潜在危险的写入升级为人工审查,这使得语义分类对于准确捕获内部人员意图变得重要。
英文摘要
Agentic AI systems with persistent memory introduce a distinct attack surface known as memory poisoning, in which adversarially crafted content is stored in long-term memory and subsequently influences future agent behavior. Such attacks can suppress security alerts, facilitate privilege escalation, alter trust relationships, or override security policies without modifying the underlying model weights or system prompts. To address this threat, we present MemSentry, a formal, configuration-driven framework that intercepts proposed persistent-memory writes and produces deterministic Accept, Review, or Quarantine decisions. MemSentry evaluates each write by jointly considering source trust, semantic risk, attack radius over a component-dependency DAG, access risk, and a signed security-state delta that captures whether an operation weakens or strengthens the system's security posture. We instantiate the protected environment using a 20-asset random dependency DAG and a 10 x 20 user access-control matrix, and evaluate the framework over 1,000 GPT-4-generated scenarios using a stratified 70/30 train/test split. Semantic classification is treated as a pluggable component rather than a primary contribution, and we compare four representative approaches: rule-based Regex, TF-IDF+SVM, SBERT+LR, and SetFit. SBERT+LR achieves the best overall performance with 91.7% accuracy and a 0.908 macro-F1 score, while all four methods detect 100% of external quarantine-class threats. For verified insiders, where source trust is maximal (T = 1), MemSentry does not automatically quarantine suspicious operations but instead escalates potentially dangerous writes for human review, making semantic classification important for accurately capturing insider intent.