AI 中文总结
该研究发布开源电信推理器TelecomGPT-R1-9B,通过两阶段后训练方案优化,在公开电信基准中表现优于同类开源模型,性能接近闭源前沿推理器。
AI 中文摘要
电信领域是基于大语言模型(LLM)推理的高杠杆领域,因为常规工程工作流程需要同时结合规范标准、运营遥测数据、厂商特定故障证据以及精确的射频(RF)/网络计算。然而,当前LLM在电信领域的集成仍受限于双向能力差距:通用推理器往往缺乏电信领域的针对性依据,而特定领域的电信LLM在结构化多步骤推理方面仍存在局限。为弥合这一差距,我们发布了TelecomGPT-R1-9B,一款在GSMA开源电信排行榜上排名领先的统一开源电信推理器。具体而言,我们整理了包含67427个样本的监督微调(SFT)语料库,围绕四个互补推理轴构建:协议、知识、建模和故障。该语料库由与各轴匹配的公开网络资源构建,并通过轴特定的思维链(CoT)生成和前缀-延续自验证进行增强。以Qwen3.5-9B为基础,我们进一步开发了两阶段后训练方案:首先,基于多教师低秩适配(LoRA)的SFT注入电信知识并诱导轴特定的推理格式;其次,通过解耦剪辑和动态采样策略优化(DAPO)稳定的组相对策略优化(GRPO),使用四个轴对齐的二元验证器奖励来优化策略。在七个公开电信基准测试中,TelecomGPT-R1-9B在开源电信LLM中排名第一,且其七轴均值可与最先进的闭源前沿推理器相媲美。
英文摘要
Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized around four complementary reasoning axes: protocol, knowledge, modeling, and fault. The corpus is built from axis-matched public web sources and enhanced through axis-specific chain-of-thought (CoT) generation and prefix-continuation self-validation. Starting from Qwen3.5-9B, we further develop a two-stage post-training recipe. First, multi-teacher low-rank adaptation (LoRA)-based SFT injects telecom knowledge and induces axis-specific reasoning formats. Second, group relative policy optimization (GRPO), stabilized by decoupled clip and dynamic sampling policy optimization (DAPO), optimizes the policy using four axis-aligned binary verifier rewards. Across seven public telecom benchmarks, TelecomGPT-R1-9B ranks first among open-source telecom LLMs and achieves a seven-axis mean comparable to state-of-the-art closed-source frontier reasoners.