AI 中文总结
提出LARGE框架结合LLM生成与迭代在环仿真,为LEO卫星网络RL路由自动设计奖励,其性能可与专家基线相当,凸显该优化方法的潜力。
AI 中文摘要
低轨(LEO)卫星网络中的路由因拓扑高度动态、网络条件具有时空特性而极具挑战性。强化学习(RL)已成为自适应路由的有前景方法,但其性能关键依赖于奖励函数设计,奖励函数需平衡吞吐量(goodput)和端到端延迟等目标。实际中,奖励设计仍是复杂的手动过程,需大量领域专业知识和反复试错。近期研究探索用大语言模型(LLM)实现自动奖励设计,但其在LEO卫星网络等高度动态系统中的应用仍基本未被探索。我们提出LARGE框架,该框架结合LLM驱动的生成与迭代的在环仿真评估,为基于RL的路由自动设计奖励。LARGE利用LLM的先验知识生成初始奖励,并通过仿真反馈迭代优化该奖励。此循环可探索多样化的奖励公式,同时使其符合网络目标。结果表明,LARGE通过反馈驱动的优化,在数次迭代内提升奖励质量。在不同主干网络下,该框架实现的性能可与专家设计的基线相当,其中表现最佳的配置的吞吐量约为基线的97%,端到端延迟略低,且无需手动进行奖励工程。这些结果表明,由LARGE实现的迭代反馈驱动过程可产生有效效果,凸显了框架驱动的LLM在环优化在动态卫星网络中基于RL的路由方面的潜力。
英文摘要
Routing in Low Earth Orbit (LEO) satellite networks is challenging due to highly dynamic topologies and spatio-temporal network conditions. Reinforcement Learning (RL) has emerged as a promising approach for adaptive routing; however, its performance critically depends on reward function design, which must balance objectives such as goodput and end-to-end delay. In practice, reward design remains a complex manual process requiring significant domain expertise and extensive trial-and-error. Recent works have explored Large Language Models (LLMs) for automated reward design, but their application to highly dynamic systems such as LEO satellite networks remains largely unexplored. We propose LARGE, a framework that automates reward design for RL-based routing by combining LLM- driven generation with iterative simulator-in-the-loop evaluation. LARGE generates an initial reward from LLM prior knowledge and iteratively refines it using simulation feedback. This loop enables exploration of diverse reward formulations while aligning them with network objectives. Results show that LARGE improves reward quality within a few iterations through feedback-driven refinement. Across different backbones, the framework achieves performance comparable to an expert-designed baseline, with the best-performing configuration reaching goodput within approximately 3% of the baseline and slightly lower end-to-end delay, without manual reward engineering. These results indicate that effectiveness emerges from the iterative feedback-driven process enabled by LARGE, highlighting the potential of framework-driven LLM-in-the-loop optimization for RL-based routing in dynamic satellite networks.
CommentsThis paper was accepted for publication at the IEEE Global Communications Conference (GLOBECOM 2026)