发表机构
Tynapse(Tynapse)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究跨大语言模型智能体的异步攻击归因问题,提出异步归因指纹向量($A^2FV$)协议,构建SCD-v1基准,实验表明$A^2FV$在活动关联上表现良好,确立跨智能体活动归因作为保护大语言模型智能体的独特评估层。
AI 中文摘要
大语言模型智能体防御通常一次评估一个会话。但在部署中,攻击可能分布在独立智能体、团队和运行时之间,每个局部防护栏只有稀疏片段。我们将跨智能体异步活动归因形式化,即在没有共享运行时状态、测试时活动标签或攻击者身份预言机的情况下,关联同一潜在对抗活动的会话。我们引入异步归因指纹向量($A^2FV$),这是一种轻量级代理端参考协议,用于根据代理可观察的工具使用、时间和提示残余对活动相似性进行成对评分。我们还构建了SCD-v1,这是一个具有良性流量、孤立攻击、多会话活动、匹配的非预言机规避和泄漏审计的受控人物匹配基准。在SCD-v1上,$A^2FV$在活动关联方面实现了0.82的成对AUC,而在相同任务下,仅对会话检测器和分块大语言模型判断进行基于分数的调整仍接近随机水平。最强的固定信号由结构和文体残余携带,而时间作为更丰富代理痕迹的诊断通道保留下来。交叉风格控制表明,该信号部分对风格敏感,但不能仅归结为风格。静态和维度感知非预言机压力测试进一步表明,在受控规避下成对可分离性仍然存在。这些结果将跨智能体活动归因确立为在野外保护大语言模型智能体的一个独特评估层。
英文摘要
LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent agents, teams, and runtimes, leaving each local guardrail with only a sparse fragment. We formalize cross-agent asynchronous campaign attribution: linking sessions from the same latent adversarial campaign without shared runtime state, test-time campaign labels, or attacker identity oracles. We introduce Asynchronous Attribution Fingerprint Vectors ($A^2FV$), a lightweight proxy-side reference protocol for scoring pairwise campaign similarity from proxy-observable tool-use, timing, and prompt residue. We also construct SCD-v1, a controlled persona-matched benchmark with benign traffic, isolated attacks, multi-session campaigns, matched non-oracle evasion, and leakage audits. On SCD-v1, $A^2FV$ achieves 0.82 pairwise AUC for campaign linking, while score-only adaptations of per-session detectors and chunked LLM judges remain near chance under the same task. The strongest fixed signal is carried by structural and stylometric residue, while timing is retained as a diagnostic channel for richer proxy traces. Crossed-style controls show that the signal is partly style-sensitive but not reducible to style alone. Static and dimension-aware non-oracle stress tests further show that pairwise separability persists under controlled evasion. These results establish cross-agent campaign attribution as a distinct evaluation layer for securing LLM agents in the wild.
Comments22 pages, 5 figures. Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at ICML 2026