AI 中文总结
针对集成TN-NTN 6G网络,提出基于Q学习的自适应信道分配框架,通过观察网络相关动态学习最优策略,设计多目标奖励函数并采用ε-贪婪策略,仿真结果显示该框架在多方面表现良好。
AI 中文摘要
本文提出了一种基于Q学习的自适应信道分配框架,其中智能体在马尔可夫决策过程(MDP)中通过观察网络负载、干扰条件和时间流量动态来学习最优策略。设计了一个多目标奖励函数来联合优化系统吞吐量、用户公平性和干扰缓解,同时采用ε-贪婪策略促进有效探索。仿真结果表明该框架稳定收敛,平均奖励为37.5,平均吞吐量为28.5Mbps,Jain公平指数为0.75,与随机分配相比干扰减少26.3%。
英文摘要
This paper proposes a self-adaptive channel assignment framework based on Q-learning, where agents learn optimal policies by observing network load, interference conditions, and temporal traffic dynamics within a Markov decision process (MDP). A multi-objective reward function is designed to jointly optimize system throughput, user fairness, and interference mitigation, while an ε-greedy strategy is employed to facilitate effective exploration. Simulation results demonstrate stable convergence, achieving an average reward of 37.5 and an average throughput of 28.5 Mbps. Moreover, the proposed approach achieves a Jain's fairness index of 0.75 and reduces interference by 26.3% compared to random allocation by adaptively responding to dynamic traffic patterns.
CommentsWe no longer stand by the results as presented and prefer to withdraw the work publicly. We apologize for any inconvenience to the community. A corrected or substantially revised version may be submitted later under a new identifier
DOI:10.1109/ICUFN69619.2026.11628659