arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TelecomGPT-R1:面向异构电信任务推理的统一后训练

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

Bohao Wang, Chenwei Wu, Hang Zou, Yu Tian, Lina Bariah, Li Wei, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Merouane Debbah

arXiv 2609.25356首次发表:更新:

发表机构

Zhejiang University; University of Michigan; Khalifa University(浙江大学; 密歇根大学; 哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对电信大模型多任务推理能力不足的问题,提出TelecomGPT-R1统一推理模型,通过轴感知数据生成、SFT和DAPO强化学习,在GSMA基准上以89.64%平均分超越GPT-5等专有模型。

AI 中文摘要

大型语言模型(LLMs)通过推理标准、网络配置、数学模型、源代码和运维日志,在自动化广泛的电信工程任务方面展现出巨大潜力。然而,现有的电信大语言模型难以在这些多样化的任务和数据类型上可靠地进行推理。通用大语言模型通常缺乏对电信特定知识的可靠基础,而电信专用模型则通常针对较窄的任务族开发,多任务性能有限。为填补这一空白,我们提出了TelecomGPT-R1,一个围绕协议、知识、建模和故障四个互补轴构建的开源统一电信推理模型系列。我们首先开发了一个轴感知的数据生成框架,将粗糙的公共电信工件精炼为经过验证的问答对和高质量的思维链(CoT)推理轨迹,从而生成包含104,880个示例的训练语料库。基于该语料库,监督微调(SFT)注入电信知识和基于证据的推理模式,以克服强化学习(RL)的冷启动障碍。然后,我们应用带有任务路由的评分奖励的动态采样策略优化(DAPO),以保持RL更新在异构电信推理任务中的信息性和稳定性。这些奖励将轴特定的CoT轨迹分解为可验证的推理单元,并将基于证据的密集过程信用与结果正确性相结合,使RL能够从可验证的电信证据中学习可泛化的问题解决行为。我们发布了TelecomGPT-R1模型和可复现的训练方案,以支持社区的进一步发展。在GSMA开放电信排行榜的七个基准上的评估显示,开源的TelecomGPT-R1-27B实现了89.64%的平均得分,超越了包括GPT-5、Claude和Gemini在内的领先专有模型。

英文摘要

Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However, existing telecom LLMs struggle to reliably reason across these diverse tasks and data types. General-purpose LLMs often lack reliable grounding in telecom-specific knowledge, while telecom-specialized models are typically developed for narrower task families and exhibit limited multi-task performance. To fill this gap, we introduce TelecomGPT-R1, a family of open source unified telecom reasoning models structured around four complementary axes: protocol, knowledge, modeling, and fault. We first develop an axis-aware data generation framework that refines coarse public telecom artifacts into verified question-answer pairs and high quality chain-of-thought (CoT) reasoning trajectories, yielding a training corpus containing 104,880 examples. Building on this corpus, supervised fine-tuning (SFT) instills telecom knowledge and evidence-grounded reasoning patterns to overcome the cold start barrier for reinforcement learning (RL). We then apply dynamic sampling policy optimization (DAPO) with task-routed rubric rewards to keep RL updates informative and stable across heterogeneous telecom reasoning tasks. These rewards decompose axis-specific CoT traces into verifiable reasoning units and combine grounded dense process credit with outcome correctness, allowing RL to learn generalizable problem solving behaviors from verifiable telecom evidence. We release the TelecomGPT-R1 models and a reproducible training recipe to support further community development. Evaluations on seven benchmarks of the GSMA Open Telco Leaderboard show that the open-source TelecomGPT-R1-27B achieves an 89.64% mean score, outperforming leading proprietary models, including GPT-5, Claude, and Gemini.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑