arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AERIS:面向多无人机集成感知与通信的离线策略改进

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

Ziyuan Wang, Yifan Sui, Wei Wei, Wenjie Xin, Zekai Zhang, Xiangwang Hou, Xiao-Ping, Zhang

arXiv 2608.25477首次发表:更新:

发表机构

Tsinghua University; Shanghai Jiao Tong University; The University of Hong Kong; The Hong Kong University of Science and Technology (Guangzhou); Toronto Metropolitan University(清华大学; 上海交通大学; 香港大学; 香港科技大学(广州); 多伦多都会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多无人机集成感知与通信的控制问题,提出离线策略改进框架AERIS及STAR-CRDT算法,在保障安全的同时显著提升ISAC性能与效率。

AI 中文摘要

无人机(UAV)赋能的集成感知与通信(ISAC)是一种颇具前景的6G范式,但动态多无人机ISAC控制需在随机移动性下协同平衡通信质量、感知可靠性与飞行安全。现有优化方法常需重复进行全局非凸求解,而在线强化学习(RL)依赖存在风险的试错飞行,可能引发感知损失或碰撞风险事件。本文提出AERIS,一种面向多无人机ISAC的离线策略改进框架。AERIS在集中式训练与分布式执行模式下从固定飞行日志中学习,各无人机基于本地历史行动,训练时则利用记录的全局信息评估团队层面的效果。我们进一步设计STAR-CRDT,一种离线多智能体强化学习算法,该算法执行支持感知的本地动作修正,并仅将可信的改进提炼至分布式执行器。我们证明了离线支持策略改进的保证。实验表明,STAR-CRDT相比最强基线,将ISAC主要目标回报提升了29.3%;同时将通信总和速率、感知通过率、感知裕度分别提升3.4%、4.8%、69.1%,碰撞风险事件减少54.2%;在由OpenStreetMap数据构建的未见过的真实道路地图上,STAR-CRDT仍取得最优回报。

英文摘要

Unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC) is a promising 6G paradigm, but dynamic multi-UAV ISAC control must jointly balance communication quality, sensing reliability, and flight safety under stochastic mobility. Existing optimization methods often require repeated global non-convex solving, while online reinforcement learning (RL) depends on risky trial-and-error flights that may cause sensing loss or collision-risk events. This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC. AERIS learns from fixed flight logs under centralized training and decentralized execution, so each UAV acts from local histories while training uses logged global information to assess team-level effects. We further design STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor. We prove an offline-support policy improvement guarantee. Experiments show that STAR-CRDT improves the main ISAC objective return by 29.3% over the strongest baseline. It further improves communication sum rate, sensing pass rate, and sensing margin by 3.4%, 4.8%, and 69.1%, while reducing collision-risk events by 54.2%. On unseen real-road maps built from OpenStreetMap data, STAR-CRDT still obtains the best return.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑