arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OTel:构建面向智能网络的领域专用电信大语言模型基础

OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks

Farbod Tavakkoli, Roderic Paulk, Jorden Terrazas, Kenneth Church, Mark Austin, Louis Powell, Gregory Diamos, Lina Bariah, Syed Ali Raza Zaidi, Maryam Hafeez, Ali Maatouk, Imtiaz Karim

arXiv 2608.15436首次发表:更新:

发表机构

AT&T Chief Data Office; GSMA; RelationalAI; Khalifa University; University of Leeds; Yale University; The University of Texas at Dallas(AT&T首席数据办公室; 全球移动通信系统协会; RelationalAI公司; 哈利法大学; 利兹大学; 耶鲁大学; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出开源电信AI资源OTel,含多类任务数据集与30个模型基线,后训练提升三类模型性能,获大量下载与媒体报道,邀请社区共建更强电信大语言模型。

AI 中文摘要

前沿AI模型发展迅速,但仍难以应对电信领域的特定任务。我们推出Open Telco(OTel),这是一个开源电信AI资源,包含用于检索、重排序、指令调优及安全/弃权(不执行)任务的衍生数据集,以及30个全参数后训练基线,覆盖嵌入、重排序和语言模型三大类。截至2026年5月3日,该资源已获得社区的大量参与:发布的模型下载量超1600万次,项目在全球获得157+家媒体报道。OTel基于现有开源电信数据集和基准构建,将有文档记录的电信数据源、预留评估分区、训练好的嵌入模型、重排序器、基于上下文的大语言模型及安全/弃权数据整合为统一资源。OTel的后训练提升了所有三类模型的性能:嵌入检索达到93.5%的NDCG@10,重排序达到0.952的MRR@10,语言模型正确率达到88.2%。我们发布OTel作为可复现的起点,邀请社区扩充数据、改进嵌入和重排序模型,并构建更强的基于上下文的电信大语言模型。

英文摘要

Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource with derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, plus 30 full-parameter post-trained baselines across embedding, reranking, and language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times, and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.5% NDCG@10, reranking reaches 0.952 MRR@10, and language-model correctness reaches 88.2%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.

CommentsAccepted at the ACM AI Leadership Summit, Breakthrough Impact Highlights Track, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑