arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一次感知,多场景服务:面向多租户ISAC网络在线感知会话整合的公共轨迹分解式约束PPO算法

Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC Networks

Dang-Dung Vu

arXiv 2608.29256首次发表:更新:

发表机构

VNU University of Engineering and Technology (VNU-UET)(越南国家大学工程技术大学(VNU-UET))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多租户ISAC网络在线感知会话整合问题,提出CT-PPO算法,通过公共轨迹奖励信用等设计,提升回报、降低感知资源成本,在多负载场景下表现优于对比算法。

AI 中文摘要

集成感知与通信(ISAC)网络可通过共享感知会话服务兼容请求,但整合过程需耦合准入控制、资源复用、配置文件选择、感知服务水平协议(SLAs)、通信服务质量(QoS)及未来承诺。我们将该问题建模为约束马尔可夫决策过程,并提出公共轨迹分解式约束近端策略优化(CT-PPO)算法。训练期间,随机策略副本共享同一原始工作负载轨迹;留一法折扣蒙特卡洛回报对比为适用的演员因子提供奖励信用,而约束信用则保持因子/前缀特异性。在5个训练种子及匹配工作负载下,CT-PPO实现最高平均宏观回报,较匹配的联合信用PPO(JC-PPO)超出0.934(95%置信区间[0.702, 1.164]),较感知SLA感知贪心算法高出1.847;与JC-PPO相比,其感知资源成本降低6.277,每创建会话的接受请求数提升0.0806。四向消融实验显示,仅分解式替代函数无明显宏观回报增益,添加公共轨迹奖励信用则产生主导性改进。无需重训练时,CT-PPO在低、标称及高到达负载下均保持回报优势,在聚类到达场景下增益最强。部署采用公共观测与硬掩码;CT-PPO的额外参数仅存在于训练侧,其演员模型规模与JC-PPO匹配,仅演员部分的CPU延迟无明显变化。

英文摘要

Integrated sensing and communication (ISAC) networks can serve compatible requests through shared sensing sessions, but consolidation couples admission, reuse, profile selection, sensing service-level agreements (SLAs), communication quality of service (QoS), and future commitments. We formulate this problem as a constrained Markov decision process and propose Common-Trace Factorized Constrained Proximal Policy Optimization (CT-PPO). During training, stochastic policy replicas share the same primitive workload trace; leave-one-out discounted Monte Carlo return contrasts provide reward credit to applicable actor factors, while constraint credit remains factor/prefix-specific. Across five training seeds and matched workloads, CT-PPO achieves the highest mean macro return, exceeding matched Joint-Credit PPO (JC-PPO) by 0.934 (95% confidence interval [0.702, 1.164]) and SLA-Aware Greedy by 1.847; versus JC-PPO, it reduces sensing-resource cost by 6.277 and raises accepted requests per created session by 0.0806. A four-way ablation shows that the factorized surrogate alone yields no detectable macro-return gain, whereas adding common-trace reward credit produces the dominant improvement. Without retraining, CT-PPO retains a return advantage at low, nominal, and high arrival loads, with the strongest gain under clustered arrivals. Deployment uses public observations and hard masks; CT-PPO's extra parameters are training-side, its actor footprint matches JC-PPO, and actor-only CPU latency is effectively unchanged.

Comments13 pages, 4 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑