arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02518cs.AIcs.CY

Magnet:通过能力累积检测跨会话的AI滥用

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

Natalie Isak, Matthew Dressman

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对AI集成体的跨会话滥用风险,提出Magnet方法,通过聚合跨会话累积的相关能力来检测此类滥用,填补了现有单会话检测的缺口。

中文摘要 AI 辅助

最强大的AI部署并非单一模型,而是由专门智能体组成的集成体,它们进行委托并协同行动。这种架构解锁了强大的新能力,但也带来了现有监控、检测和缓解框架未设计应对的风险。大多数最先进的AI滥用检测文献聚焦于单轮或多轮(单会话)威胁模型,这留下了关键缺口:攻击者可将有害目标分解为看似无害的单元,并在孤立的智能体会话中执行每个单元。智能体在对话间无状态,但攻击者并非如此,这种不对称性催生了能有效逃避检测的跨会话轨迹。我们的贡献有两点:第一,我们展示了跨会话目标分解作为一种逃避技术,表明它可能比等效的单会话或多轮攻击引发更有害的能力。此处的能力指目标某一步产生的产物,由交互产生的内容(模型响应和工具调用结果)证明,且可与其他地方累积的能力组合成有害整体。第二,我们提出了Magnet:一种高效且鲁棒的检测方法,该方法对随时间及跨智能体对话累积的相关能力进行建模,在更高层级的关联器(此处为用户ID)处聚合,而非按会话状态聚合。主要挑战是组装Magnet所推理的证据束:有罪的产物可能是散落在大量良性会话中的“针头”,这些会话单独无害,仅在收集后才具危险性。Magnet不像逐根查找干草中的针(即按会话检查),而是如其名所示:它将相关的“针头”从干草中吸引出来,跨会话和时间收集到紧凑的证据束中,供检测器采取行动。

英文摘要

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a critical gap: an attacker can decompose a harmful goal into innocuous-looking units and execute each in isolated agentic sessions. The agent is stateless between conversations, but the attacker is not. This asymmetry allows for cross-session trajectories that are effective at evading detection. Our contributions are twofold. First, we demonstrate cross-session goal decomposition as an evasion technique, showing it may elicit more harmful capability than equivalent single-session or multi-turn attacks. By capability we mean an artifact produced at one step of an objective, evidenced by what an interaction produced (model responses and tool-call results), and composable with capabilities accrued elsewhere into a harmful whole. Second, we propose Magnet: an efficient and robust detection approach that models relevant capabilities accrued over time and across agentic conversations, aggregated at a higher-level correlator (in this case, a user ID) rather than per-conversation state. The main challenge is assembling the evidence bundle Magnet reasons over. The incriminating artifacts may be needles scattered through a haystack of benign sessions that are individually harmless, dangerous only once collected. Rather than searching the haystack straw-by-straw (i.e. per-session inspection), Magnet does what its name implies: it attracts the relevant needles out of the hay, across sessions and across time, into a compact evidence bundle a detector can act on.

↑