H2LooP 电信模型 v1:从电信理解到自主问题与 PR 解决
H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution
浏览论文内容
中文总结 AI 辅助
本文提出 H2LooP 电信模型 v1,通过领域微调实现电信理解与自主代码生成,在 OT-Lite 等基准上超越闭源模型,且保持通用能力无遗忘。
中文摘要 AI 辅助
我们提出了 H2LooP 电信模型 v1,这是一个为电信行业微调的专业领域大语言模型。我们发布了两个领域适配的模型变体,服务于互补的用例:一个面向理解的变体,用于电信领域问答和推理;一个智能体变体,用于自主电信代码生成、拉取请求解决以及生产仓库上的代码提交。H2LooP 电信在 GSMA 开放电信精简版(OT-Lite)基准和一个专有电信代码生成基准上取得了强劲结果,在独立排行榜评估中超越了 GPT-5 和 Claude Opus 等前沿闭源模型,同时保持了通用能力。理解变体在 OT-Lite Pass@3 上实现了 81.8% 的加权平均分,并且独立地在官方社区运营的开放电信 AI 排行榜*上以仅 31B 参数排名第 5,领先于包括 Claude Opus 4.6、GPT-5、Gemini 3 Flash、Grok-4-fast 和 Kimi K2.5 在内的前沿闭源系统。我们的智能体变体在电信代码生成上相对于基础模型获得了 +8.8% 的 AST 相似度和 +20.0% 的位置 IoU 的相对改进,同时保持了相同的 MMLU(74.0%)和 BFCL v3 多轮函数调用(79.0%)性能,表明零灾难性遗忘。在精选的电信语料库(涵盖 3GPP 标准、O-RAN 规范、网络遥测和真实仓库提交)上的领域专业化,相对于同等规模的通用模型产生了显著改进,在领域特定评估上接近前沿闭源模型,并由我们的官方排行榜地位独立证实。
英文摘要
We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the telecommunications industry. We release two domain-adapted model variants serving complementary use cases: a comprehension-focused variant for telecom domain question answering and reasoning, and an agentic variant for autonomous telecom code generation, pull request resolution, and code commits on production repositories. H2LooP Telecom achieves strong results on the GSMA Open Telecom Lite (OT-Lite) benchmark and a proprietary telecom code generation benchmark, outperforming frontier closed-source models such as GPT-5 and Claude Opus on independent leaderboard evaluation, while preserving general-purpose capabilities. The Comprehension variant achieves 81.8% weighted average on OT-Lite Pass@3, and, independently, ranks 5th overall on the official community-run Open Telco AI Leaderboard* at only 31B parameters-ahead of frontier closed-source systems including Claude Opus 4.6, GPT-5, Gemini 3 Flash, Grok-4-fast, and Kimi K2.5. Our agentic variant obtains a relative improvement of +8.8% in AST Similarity and +20.0% in Location IoU over the base model on telecom code generation, while maintaining identical MMLU (74.0%) and BFCL v3 multi-turn function calling (79.0%) performance, indicating zero catastrophic forgetting. Domain specialization on curated telecom corpora, spanning 3GPP standards, O-RAN specifications, network telemetry, and real repository commits, yields substantial improvements over general-purpose models of equivalent scale, approaches frontier closed-source models on domain-specific evaluation, and is independently corroborated by our official leaderboard standing.
发表机构
- H2LooP.ai
机构由 AI 辅助整理,请以论文原文为准。