ReForge:通过经过验证的大语言模型编辑让ABR算法永不落后
ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
- The University of Electro-Communications(电气通信大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出ReForge框架,将LLM纳入循环以对ABR算法的模糊规则进行验证性小编辑,使其适配持续变化的网络场景,在9个真实网络系列上大幅提升QoE性能。
AI中文摘要:
为单一网络场景设计一个ABR算法需要工程师花费数月时间,而如今大语言模型可在数小时内完成这项工作,性能与人工构建的设计相当甚至更优。但无论哪种方式,该设计仅适配其诞生时可见的网络环境,对之后出现的新网络环境失效。我们提出疑问:ABR算法能否跟上网络环境的变化,在每个新场景出现时于数分钟内完成重新设计,且每次变更都要确保对已服务的所有场景无害。本研究提出ReForge,这是一个适配持续变化场景的持续启发式学习框架,该框架将大语言模型(LLM)纳入循环运行。每一轮中,LLM会读取当前设计的不足并提出一个小的编辑方案,通过对迄今为止服务过的所有网络进行重放来决定是否采纳该编辑。具体而言,其编辑对象是一页模糊规则,该规则将每个决策路由到一个冻结的预训练策略池中。LLM仅通过测量结果生成初始规则页,之后自行不断优化。每一轮中,LLM读取当前规则的不足并提出一个小的编辑方案,再通过对迄今为止服务过的所有网络进行重放来决定该编辑是否有效。我们在9个依次出现的真实网络系列(3G、4G、5G)上评估ReForge,每次新网络系列到来时,通过数次编辑将平均QoE从1.23提升至1.74,超过最佳单一策略的1.66,达到神谕(oracle)性能的94%,甚至能修复循环从未见过的网络系列,其中一个网络系列的QoE从0.30升至0.80。所有代码、数据和实验记录将在清理后开源。
英文摘要:
Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching or beating hand-built designs. But either way, the design fits only the world visible at its birth, and fails on the world that arrives after. We ask whether an ABR algorithm can keep pace with the world, redesigned in minutes as each scenario arrives, with every change proven harmless to every scenario already served. In this work, we propose ReForge, a continual heuristic learning framework that adapts to continuously changing scenarios. ReForge runs that routine with a large language model (LLM) in the loop. Each round the LLM reads where the current design falls short and proposes one small edit, and a replay over every network served so far decides. Specifically, what it edits is a single page of fuzzy rules that routes every decision to one of a frozen pool of pre-trained policies. The LLM writes the first page from measurements alone, then keeps improving it on its own. Each round it reads where the current rules fall short and proposes one small edit, and a replay over every network served so far decides whether the edit lands. We evaluate ReForge on nine real-world network families arriving one at a time as 3G, 4G, then 5G. A few edits per arrival lift mean QoE from 1.23 to 1.74, past the best single policy at 1.66 and to 94\% of an oracle, and even repair families the loop never saw, one rising from 0.30 to 0.80. All code, data, and experiment records will be open-sourced upon cleanup.