Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
多智能体人工智能控制:分布式攻击阻碍实例级监测
机构 * MATS Research(MATS研究所)
AI总结 研究多智能体人工智能控制中的分布式攻击,开发FakeLab代码库,评估单智能体监测应对分布式攻击的效果,发现碎片化效应,指出其与代码比例无关,受模型能力影响,还揭示显式规划器及不同强度监测的作用。
Comments Submitted to NeurIPS; 81 pages; 32 figures and 24 tables