arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00590cs.LG

CRAFT:原生AI 6G无线接入网中预可解释性的微调

CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN

  • NextG Wireless Lab(下一代无线实验室)
  • NCSU(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Pranshav Gajjar, Vijay K Shah

AI总结:

本文针对6G RAN中电信LLM的预可解释性冷启动障碍,提出CRAFT方法,通过生成验证数据集微调SLM,在低能耗下实现高准确率与F1值,为可审计AI部署提供可行方案。

AI中文摘要:

下一代移动网络被设想为完全原生AI的网络,其中AI-RAN架构嵌入小型语言模型(SLM)以对实时遥测数据进行推理。电信领域大型语言模型(LLM)的最先进训练范式,例如RANSTRUCT风格的、基于精心整理的指令数据的监督微调(SFT),仅局限于事后解释。这种解释若能生成,也是在决策之后或独立于决策产生,导致决策过程不可审计。而预推理(即先于输出标签生成因果推理轨迹)更受青睐,更广泛的LLM推理文献已通过Group Relative Policy Optimization(GRPO)等强化学习方法在这方面取得了切实进展。我们发现,将该方案迁移到电信场景会遭遇冷启动障碍:SLM要么学会输出所需格式,要么学会预测标签,但很少能同时做到两者。我们识别出该障碍并提出CRAFT,即Cold-start Reasoning Alignment via Fine-Tuning(通过微调实现冷启动推理对齐),这是一种以数据为中心的方法,可自主生成经过验证的(输入、轨迹、标签)三元组数据集。CRAFT使用低秩适配(LoRA)在该验证数据上对SLM进行微调,所需计算量和时钟时间远少于基于GRPO的方法。在TRACTOR和IC xApp电信数据集上,CRAFT在无解析失败的情况下,准确率和F1值分别达到86.5%和94.6%,而直接GRPO及SFT+GRPO方法的F1值均未超过28%和53.5%,且存在多次解析失败。我们进一步证明,以CRAFT初始化的策略可为后续GRPO微调提供稳健基础,在不同奖励函数下,其性能保持一致且无解析失败。最后,我们展示CRAFT的能耗比基于GRPO的基准方法低59%,是6G RAN中部署可审计AI的可持续路径。

英文摘要:

The next generation of mobile networks is envisioned as fully AI-native, with AI-RAN architectures embedding small language models (SLMs) to perform reasoning over real-time telemetry. The state-of-the-art training paradigms for telecom LLMs, exemplified by RANSTRUCT-style supervised fine-tuning (SFT) on curated instruction data, are limited to post hoc rationalization. Here, the explanations, when produced at all, are generated after or independently of the decision, leaving the decision process unauditable. Pre-hoc reasoning, where a causal reasoning trace is produced before the output label, is preferable, and the broader LLM reasoning literature has made real progress toward it via RL methods such as Group Relative Policy Optimization (GRPO). Here we observe that transplanting this recipe into the telecom setting runs into a cold-start barrier: SLMs either learn to output the desired format or learn to predict the label, but rarely both. We identify this barrier and propose CRAFT, which stands for Cold-start Reasoning Alignment via Fine-Tuning, a data-centric method to autonomously generate a verified dataset of (input, trace, label) triplets. CRAFT fine-tunes SLMs on this verified data using low-rank adaptation (LoRA), requiring substantially less compute and wall-clock time than GRPO-based methods. On the TRACTOR and IC xApp telecom datasets, CRAFT achieves up to 86.5% and 94.6% for accuracy and F1 with no parse failures, while direct GRPO and SFT+GRPO fail to exceed 28% and 53.5% F1 with multiple parse failures. We further show that CRAFT-initialized policies serve as a robust foundation for subsequent GRPO fine-tuning, as under diverse reward functions the performance remains consistent with no parse failures. Finally, we demonstrate that CRAFT consumes 59% less energy than GRPO-based baselines, making it a sustainable path to deployable, auditable AI in 6G RAN.

↑