arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11121eess.SYcs.SY

量化基于强化学习的无人机在毫米波和亚太赫兹频段部署的现实差距

Quantifying the Reality Gap for RL-Based UAV Placement at mmWave and Sub-THz

Abdullateef Almohamad, Mostafa Ibrahim, Sabit Ekin, Khalid Qaraqe

首次发表
浏览论文内容

中文总结 AI 辅助

本研究量化了毫米波和亚太赫兹频段无人机部署中强化学习策略的仿真到现实差距,通过三种信道模型对比,发现偏差主要源于蒙特卡洛欠采样和大气吸收模型差异,但策略仍接近最优。

中文摘要 AI 辅助

用于毫米波和亚太赫兹网络中无人机(UAV)部署的强化学习(RL)策略通常在简化的解析信道上进行训练。我们在卡塔尔多哈的真实城市地图上,在载波频率{28, 140, 183, 300} GHz和高度{50, 75, 100, 125} m下,量化了由此产生的仿真到现实差距,评估了三种信道流水线:解析模型(FSPL + 大气吸收 + 长方体视距)、Sionna RT中带有ITU-R P.676-13吸收的完整蒙特卡洛射线追踪,以及一种在闭式路径增益表达式下重用Sionna网格的确定性视距混合模型。我们通过四个指标,即偏差、均方根误差、Jensen-Shannon散度和最优部署位移,在空间信噪比分布上形式化了这一差距。出现了三个发现:在28/140 GHz下,Sionna明显的-5.6/-4.8 dB偏差中约70%是蒙特卡洛欠采样造成的,在缓解后缩小到-1.7/-1.5 dB;在183 GHz下,-9.2 dB的残余偏差隔离了大气吸收/ITU-R P.676线形不一致的问题;在300 GHz下,随机射线追踪器与解析模型仅偶然一致,确定性视距流水线暴露了+3.8 dB的结构性偏移。在所有载波频率下,解析训练策略的线性域遗憾值保持≥0.93,表明实际接近最优,但存在随载波变化的信噪比偏差,值得明确报告。

英文摘要

Reinforcement learning (RL) policies for unmanned aerial vehicle (UAV) placement in mmWave and sub-terahertz networks are typically trained on simplified analytical channels. We quantify the resulting sim-to-real gap on a real urban map of Doha, Qatar, at carriers {28, 140, 183, 300} GHz and altitudes {50, 75, 100, 125} m, evaluating three channel pipelines: an analytical model (FSPL + atmospheric absorption + cuboid LoS), full Monte-Carlo ray tracing in Sionna RT with ITU-R P.676-13 absorption, and a deterministic-LoS hybrid that reuses Sionna's mesh under a closed-form path-gain expression. We formalize the gap on the spatial SNR distribution via four metrics, namely bias, RMSE, Jensen-Shannon divergence, and optimum-deployment displacement. Three findings emerge: at 28/140 GHz, $\sim$70% of the apparent -5.6/-4.8 dB Sionna bias is Monte-Carlo undersampling and shrinks to -1.7/-1.5 dB after mitigation; at 183 GHz a -9.2 dB residual isolates the atmospheric absorption / ITU-R P.676 line-shape disagreement; at 300 GHz the stochastic ray tracer agrees with the analytical model only coincidentally, with a +3.8 dB structural offset exposed by the deterministic-LoS pipeline. Across all carriers the linear-domain regret of the analytical-trained policy stays $\geq$ 0.93, indicating practical near-optimality but with a carrier-resolved SNR bias that warrants explicit reporting.

发表机构

  • Texas A&M University(德克萨斯农工大学)
  • Hamad Bin Khalifa University(哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑