arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26518eess.SYcs.SY

对抗性无人机巡逻中作为可隐藏状态的能量:公式化、能量-安全阈值及自我博弈的局限性

Energy as a Concealable State in Adversarial UAV Patrolling: Formulation, an Energy-Security Threshold, and the Limits of Self-Play

Sai Krishna Reddy Mareddy

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对图上受能量约束的对抗性无人机巡逻问题,构建零和部分可观察随机博弈,发现能量预算决定防御能力,现有学习算法无法达稳定均衡,攻击者会集中攻击充电窗口,防御者可学习欺骗性充电时机。

中文摘要 AI 辅助

我们研究图上受能量约束的对抗性巡逻问题,其中电池容量有限的无人机需防御一组高价值目标,抵御选择攻击时机与地点的策略型攻击者。与现有对抗性巡逻不同,巡逻者必须定期返回基地充电;与现有能量感知型巡逻不同,巡逻者面临自利型攻击者。我们的核心发现是剩余能量是隐藏状态:攻击者无法直接观测电池电量,但可观测巡逻者轨迹,推断何时即将进行充电行程(即易受攻击窗口)。我们将该交互形式化为零和部分可观察随机博弈,并报告求解器层面的负面结果:独立深度Q学习与神经虚构自我博弈均无法在该规模下达到稳定均衡;两者均会短暂提升性能后崩溃,协同训练的挫败率在训练过程中从约0.25降至约0.09。利用不依赖学习动态的结构分析,我们表明可实现的安全性随能量预算单调上升,在阈值以下为0,预算充足时升至约0.7,确立能量预算为防御能力的主要决定因素。我们阐述该模型旨在回答的问题:具备推理能力的攻击者是否会将成功攻击集中在充电窗口,防御者是否能学习欺骗性充电时机以关闭该窗口。

英文摘要

We study energy-constrained adversarial patrolling on a graph, in which a battery-limited UAV defends a cluster of high-value targets against a strategic attacker who chooses when and where to strike. Unlike prior adversarial patrolling, the patroller must periodically return to a base to recharge; unlike prior energy-aware patrolling, it faces a self-interested adversary. Our central observation is that the remaining energy is a hidden state: the attacker never observes the battery directly, but observes the patroller's trajectory and can infer when a recharge excursion, and thus a vulnerability window, is imminent. We formalize the interaction as a zero-sum partially observable stochastic game and report a negative result on the solver side: neither independent deep Q-learning nor Neural Fictitious Self-Play reaches a stable equilibrium at this scale; each improves transiently and then collapses, with the co-trained thwart rate falling from about 0.25 to about 0.09 over training. Using a structural analysis independent of the learning dynamics, we show that achievable security rises monotonically with the energy budget, from zero below a threshold to about 0.7 when the budget is ample, establishing the energy budget as the primary determinant of defensibility. We set out the program the model is built to answer: whether an inference-capable attacker concentrates its successful strikes in the recharge window, and whether the defender can learn deceptive recharge timing to keep that window closed.

↑