发表机构
University of Klagenfurt(克拉根福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体群体分布式集体侦察,提出SwarmReconGuard黑盒基准,比较多种检测器,发现高斯似然比检测在已知攻击上表现优异但泛化差,混合CUSUM在规模上更有效,揭示策略泛化差距。
AI 中文摘要
自主和智能体客户端可以将侦察分布到多个身份上,使得每个请求保持有效、低速率且看似良性,而群体整体上却获取了广泛的系统知识。我们将这一威胁形式化为分布式集体侦察(DCR),并提出了SwarmReconGuard,一个可复现的黑盒基准测试,其中防御者仅观察服务边界遥测。该Docker隔离研究评估了10至10,000个虚拟身份上的11种良性行为和攻击行为,包含440次测试运行和3,666,300个请求,并保证了完整的遥测完整性。我们比较了基于语义、高斯、条件、图、核、混合和CUSUM的检测器。高斯似然比检测在已知攻击上实现了100%的检测率,观察到的误报率为0%,但在未见策略上仅为3%。CUSUM在1.25%的误报率下实现了36.1%的总体检测率,而混合CUSUM在10,000个身份下达到了85.7%的检测率,观察到的误报率为0%。结果揭示了重大的策略泛化差距,并激励了暴露感知和规模感知的防御措施。
英文摘要
Autonomous and agentic clients can distribute reconnaissance across many identities so that each request remains valid, low-rate, and benign-looking while the population collectively acquires broad system knowledge. We formalize this threat as Distributed Collective Reconnaissance (DCR) and present SwarmReconGuard, a reproducible black-box benchmark in which the defender observes only service-boundary telemetry. The Docker-isolated study evaluates 11 benign and attack behaviors across 10-10,000 virtual identities, comprising 440 test runs and 3,666,300 requests, with complete telemetry integrity. We compare semantic, Gaussian, conditional, graph, kernel, hybrid, and CUSUM-based detectors. Gaussian likelihood-ratio detection achieves 100$\%$ detection with 0$\%$ observed false positives on known attacks but only 3$\%$ on unseen policies. CUSUM yields 36.1$\%$ overall detection at 1.25$\%$ false positives, while hybrid CUSUM reaches 85.7$\%$ detection with 0$\%$ observed false positives at 10,000 identities. Results expose a major policy-generalization gap and motivate exposure-aware, scale-aware defenses.