arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07766cs.AI

OTel:开放电信人工智能数据集、基准与模型

OTel: Open Telco AI Datasets, Benchmarks, and Models

Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali… 展开作者

Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying, Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani, Somanshu Singla, Adarsh Chaluvaraju

首次发表
浏览论文内容

中文总结 AI 辅助

OTel是开放的电信AI资源,提供数据集、30个后训练基线模型和基准,显著提升检索、重排序及语言模型性能,并获广泛社区采用。

中文摘要 AI 辅助

我们推出开放电信(OTel),一个开放的电信人工智能资源,发布了用于检索、重排序、指令微调以及安全/弃权(不执行)的衍生电信数据集,同时提供30个全参数后训练基线模型,涵盖10个嵌入模型、3个重排序器和17个语言模型。社区已对该资源进行了大量参与:截至2026年5月3日,已发布模型下载量超过1600万次,项目在全球获得157+篇媒体报道。基于先前的开放电信数据集和基准,OTel在一个统一资源中提供了有文档记录的电信数据源、留出评估分区、训练好的嵌入模型、重排序器、基于上下文的LLM以及安全/弃权(不执行)数据。每个基线从开放权重模型开始,使用开放训练方案在OTel衍生数据上进行后训练,然后在留出的OTel评估分区上进行评估。OTel后训练在所有三个模型家族中均提升了性能:嵌入检索达到93.1%的NDCG@10,重排序达到0.947的MRR@10,语言模型正确率达到87.8%。我们将OTel作为可复现的起点发布,并邀请社区扩展数据、改进嵌入和重排序模型,以及构建更强的基于上下文的电信LLM。

英文摘要

We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.

发表机构

  • AT&T Chief Data Office(AT&T首席数据办公室)
  • RelationalAI
  • MLCommons
  • The University of Texas at Dallas(德克萨斯大学达拉斯分校)
  • Purdue University(普渡大学)
  • Khalifa University(哈利法大学)
  • University of Leeds(利兹大学)
  • Yale University(耶鲁大学)
  • Mantis NLP
  • GSMA(全球移动通信系统协会)
  • Essential AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑