arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSWM:面向需求驱动无人机基站重定位的分解时空世界模型

DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning

Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen, Xianbin Cao, Xinmei Qiang, Dapeng O. Wu, Tony Q. S. Quek

arXiv 2609.36845首次发表:更新:

发表机构

Beihang University; Peng Cheng Laboratory; Hangzhou International Innovation Institute, Beihang University; City University of Hong Kong; Hong Kong Generative AI Research and Development Center(北京航空航天大学; 鹏城实验室; 北京航空航天大学杭州国际创新研究院; 香港城市大学; 香港生成式人工智能研究发展中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DSWM提出分解时空世界模型,通过潜在决策时规划与想象展开,在三个真实数据集上以0.889-0.908的服务比率超越14种基线,实现需求驱动的无人机基站高效重定位。

AI 中文摘要

无人机基站(UAV-BSs)有望覆盖随时间和空间变化的流量需求,然而大多数重定位方案要么在每个时隙重新求解优化问题,要么学习无显式需求模型的反应式策略。我们将需求驱动的车队重定位视为潜在空间中的决策时规划问题,并提出DSWM,一种分解的时空世界模型:一个智能体控制器,通过滚动观测窗口感知需求场,在潜在循环状态中保留操作上下文,在不确定性惩罚下通过想象展开推理候选运动,并通过重新规划的首个动作协调车队。DSWM学习一个循环状态空间模型,该模型由基于指数移动平均(EMA)的潜在预测目标(带有方差正则化)塑造。它附加了一个可微分的服务模拟器,在潜在展开中重放关联、概率视距信道和香农速率链。规划使用交叉熵方法,其想象需求锚定在当前观测窗口上,混合系数$\ ho=0.95$。在统一流程上,基于三个真实数据集(Milan CDR(通话详细记录)、Shanghai Telecom、YJMob100K)和14种方法(包括五个复现的IEEE基线),DSWM在非消融配置中于每个数据集上均排名第一,工作日服务比率分别为0.889、0.908和0.898。在Milan上,它比最强的非学习基线(Greedy,0.780)提高了0.109,这一差距来自决策时对观测的使用而非预测精度。

英文摘要

Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient $ρ=0.95$. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.

Comments13 pages, 13 figures, 4 tables, 2 algorithms. Submitted to IEEE Journal on Selected Areas in Communications (Special Issue on Agentic AI for Intelligent Networks)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑