发表机构
The University of Hong Kong; Southeast University; Tencent Hunyuan Frontier Lab(香港大学; 东南大学; 腾讯混元前沿实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多任务无线网络优化中跨任务知识复用困难的问题,提出FARM奖励空间迁移框架,通过智能体奖励模型学习任务条件奖励先验并冻结复用,在MEC任务上相比单任务SAC获得29.8%的平均后期增益。
AI 中文摘要
未来无线网络需要学习智能体适应异构的信道条件、流量模式、服务质量(QoS)需求、目标以及运营约束。跨此类任务复用决策知识具有挑战性,因为传统的多任务和迁移强化学习方法主要共享或迁移策略,将可迁移知识与任务相关的动作映射耦合在一起。本文提出FARM(面向多任务无线网络优化的基础智能体奖励模型),这是一种奖励空间迁移框架,将跨任务知识复用从策略空间转向轨迹级决策评估。FARM引入了一个智能体奖励模型(ARM),该模型从异构源任务轨迹中学习任务条件的奖励先验,并为特定任务的策略优化提供辅助指导。在第一阶段,ARM联合建模任务条件、时间轨迹依赖性和目标相关的奖励结构,同时每个源任务保留自己的控制器。在第二阶段,学习到的奖励先验被冻结并复用,以指导目标特定控制器对先前未见任务的适应,而不迁移源任务策略。在异构多接入边缘计算(MEC)任务上的实验表明,对于未见过的速率-延迟目标,FARM相比单任务SAC实现了29.8%的平均后期增益,而CRA迁移为16.1%,在中等分布外(OOD)的FAR-M案例中达到了46.2%的增益。进一步分析表明,Mamba和Transformer轨迹编码器均支持奖励空间迁移,而随着更长历史依赖的引入,Mamba提供了更好的鲁棒性。
英文摘要
Future wireless networks require learning agents to adapt across heterogeneous channel conditions, traffic patterns, quality-of-service (QoS) requirements, objectives, and operational constraints. Reusing decision knowledge across such tasks is challenging because conventional multi-task and transfer reinforcement learning methods primarily share or transfer policies, coupling transferable knowledge with task-dependent action mappings. This paper proposes FARM (Fundamental Agentic Reward Model for Multi-task Wireless Network Optimization), a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation. FARM introduces an Agentic Reward Model (ARM) that learns a task-conditioned reward prior from heterogeneous source-task trajectories and provides auxiliary guidance for task-specific policy optimization. In Stage I, ARM jointly models task conditions, temporal trajectory dependencies, and objective-dependent reward structures while each source task retains its own controller. In Stage II, the learned reward prior is frozen and reused to guide the adaptation of a target-specific controller for previously unseen tasks, without transferring source-task policies. Experiments on heterogeneous multi-access edge computing (MEC) tasks show that FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, compared with 16.1% for CRA Transfer, and reaches a 46.2% gain on the moderate-OOD FAR-M case. Further analysis shows that both Mamba and Transformer trajectory encoders support Reward-Space Transfer, while Mamba provides improved robustness as longer history dependencies are introduced.