发表机构
Keele University(基尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对使用工具的LLM智能体群体的自适应隐式合谋,构建结合多类方法的黑盒隐写分析检测器,开展红蓝对抗实验,绘制检测容量边界并发现新的规避手段。
AI 中文摘要
基于大语言模型(LLM)构建的使用工具的智能体,越来越多地由多个主体而非单个操作员在共享基础设施上并行部署。这产生了单智能体安全防护措施未能覆盖的群体层面风险:少数智能体可悄悄协调,操纵市场、在评审流程中互相抬升、或同步窃取数据,而每个智能体看起来都完全合规。难题在于,运行这些智能体的组织无法查看彼此模型的内部情况,因此任何现实的检测器必须仅基于行为工作:黑盒、仅跟踪行为、且通常只有部分可见性。我们将隐式协调视为信息隐藏问题,构建了一种黑盒隐写分析检测器,它结合了跨运行互信息估计、置换检验、分布偏移统计以及时间和工具调用侧信道,所有这些都针对固定的误报预算进行了校准。我们的核心举措是不再针对单一固定编码进行测试:我们让检测器对抗不断重写编码以逃避检测器已掌握知识的对手,并在使用工具、具备记忆能力的环境而非玩具游戏中开展这种红蓝对抗。随后,容量理论为我们指明了预期结果,即检测容量边界,低于该边界的隐式比特率下,黑盒检测的表现可证明与随机猜测无异。我们设计了一项实验来绘制该边界,报告了等待测量的明确标记的占位结果,并指出了当前方法大多未发现的一种实用规避手段:在多个会话中分散负载。
英文摘要
Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrastructure. This creates a population-level risk that single-agent safeguards miss: a handful of agents can quietly coordinate, rigging a market, boosting one another in a review process, or timing a joint data grab, while each one looks perfectly well-behaved. The difficulty is that the organisations running these agents cannot see inside one another's models, so any realistic detector must work from behaviour alone: black-box, trace-only, and often with only partial visibility. We treat covert coordination as an information-hiding problem and build a black-box steganalysis detector that combines cross-run mutual-information estimation, permutation tests, distributional-shift statistics, and timing and tool-call side channels, all calibrated to a fixed false-positive budget. Our central move is to stop testing against a single fixed code: we pit the detector against an adversary that continually rewrites its encoding to slip past whatever the detector has learned, and we run this red-versus-blue contest in tool-using, memory-carrying environments rather than toy games. Capacity theory then tells us what to expect, a detection-capacity frontier, a covert bit-rate below which black-box detection is provably no better than chance. We set out an experiment to map this frontier, report clearly labelled placeholder results pending measurement, and flag a practical evasion, spreading a payload across sessions, that current methods largely miss.
Comments10 pages