arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06966cs.LGcs.CLcs.CR

MOLE:检测AI智能体中的内部威胁

MOLE: Detecting Insider Threats in AI Agents

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Aashiq Muhamed, Virginia Smith

中文总结 AI 辅助

MOLE是一个包含150个AI账户、30天、12种威胁的开放基准,用于检测AI智能体中的内部威胁,支持监控器比较与开发,显著提升检测性能。

中文摘要 AI 辅助

模型错位、提示注入或操作员误用可能导致操作前沿实验室账户的AI智能体窃取模型权重、投毒训练数据或削弱发布门控。现有基准测试并未检验防御者能否在有限的审查预算下,于日常工作中检测到此类活动。我们引入了MOLE,一个开放的基准测试,包含150个AI运营账户,在30个工作日内共享9个有状态服务,涵盖12种威胁和来自四个模型的8个语料库,总计约200亿个令牌。在39个智能体模型中,72%完成了大多数指定的有害目标,且智能体的拒绝行为并不能预测任务完成情况。MOLE支持在语料库生成器、可观测性级别和威胁类型之间比较40种监控器;即使在我们单日审计事件比较中表现最佳的监控器,也遗漏了近一半的已完成危害。MOLE还支持监控器开发:基准引导的搜索将中等监控器提升了49-64%,而选择性使用更强的监控器,在可比建模成本下,相比应用于每个账户日,将预算AUC提高了10%。

英文摘要

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.

↑