arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

太阳能开放2技术报告

Solar Open 2 Technical Report

Sungrae Park, Sanghoon Kim, Gyoungjin Gim, Jungho Cho, Hyunwoong Ko, Minbyul Jeong, Minjeong Kim, Keunwoo Choi, Chaehun Shin, Chanwoong Yoon, Dongjun Kim, Eunwon Kim, Gyungin Shin, Hyeonju Lee, Hyungkyu Kang, Inseo Song, Jisu Bae, Jiyoon Han, Jiyun Lee, Joonkee Kim, Junyeop Lee, Mikyoung Cha, Sangwon Yu, Sehwan Joo, Seokyoon Kang, Seonghoon Yang, Seung Shin, Seunghyun Lee, Seungseop Lim, Seungyoun Shin, Sukyung Lee, Taegyeong Eo, Taehwan Oh, Taewhoo Lee, Wonho Song, Wonjun Oh, Wonseok Hwang, Yunsu Kim, Yura Shim, Hwalsuk Lee, Sunghun Kim, Du-Seong Chang, Kyunghyun Cho, Seungju Han, Yejin Choi, Junsuk Choe, Hwaran Lee, Minjeong Ban, Yun Taewon, Hwanjun Song, Jae-Gil Lee, KyungTae Lim, Alice Oh

arXiv 2607.20062首次发表:更新:

AI 中文总结

本文介绍为长期代理任务构建的太阳能开放2语言模型。通过混合注意力堆栈等技术扩大规模,利用更强起点和高价值数据实现高效训练,经多教师策略蒸馏构建代理技能,在多项基准测试中表现出色,展示了其在语言模型领域的优势。

AI 中文摘要

我们展示了太阳能开放2,这是一个为长期代理任务构建的250B - A15B专家混合语言模型,从太阳能开放1(太阳能开放100B)扩展而来。为在单个上下文中保存整个代理轨迹,太阳能开放2通过混合注意力堆栈达到1M令牌窗口,该堆栈在每三个线性注意力层中交错一个softmax层,不使用位置编码和扩展到负特征值的门控增量规则。为在固定计算预算下进行这种规模的训练,我们通过更强的起点和更高价值的数据使训练高效。起点方面,从太阳能开放1初始化,转移架构变化后幸存的56.9亿参数共享框架并通过全预训练学习其他内容。数据方面,按每个令牌的价值精心挑选,通过质量和稀有度感知的数据策划和混合比例优化将20T池细化为10T混合物,在相同令牌预算下优于太阳能开放1的方法。为构建其代理技能,我们在专门场景中训练12个领域专家,然后通过多教师策略蒸馏(MOPD)将它们整合到单个模型中。在英语基准测试中,太阳能开放2在MMLU - Pro、LiveCodeBench和APEX - Agents代理套件上领先,在其他地方与最强模型(DeepSeek - V4 - Flash和MiMo - V2.5)竞争。在韩语基准测试中,太阳能开放2在所有比较模型中平均得分最高,包括快速层级的封闭API,在内部韩语办公代理基准测试Ko - GDPval上,它以不到DeepSeek - V4 - Pro(1.6T)六分之一的规模与之竞争。

英文摘要

We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, we make training efficient in two ways: a stronger starting point, and higher-value data. For the starting point, we initialize Solar Open 2 from Solar Open 1, transferring the 5.69B-parameter shared skeleton that survives the architectural change and learning everything else through full pre-training. For the data, we curate for value per token: quality- and rarity-aware data curation and mixture-ratio optimization refine a 20T pool into a 10T mixture that, at equal token budget, outperforms the Solar Open 1 recipe. To build its agent skills, we train twelve domain specialists across purpose-built scenarios, then consolidate them into a single model by Multi-teacher On-Policy Distillation (MOPD). Against comparably sized open-weight models on English benchmarks, Solar Open 2 leads on MMLU-Pro, LiveCodeBench, and the APEX-Agents agentic suite, and stays competitive with the strongest (DeepSeek-V4-Flash and MiMo-V2.5) elsewhere. On Korean benchmarks, Solar Open 2 records the highest average of any model compared, including fast-tier closed APIs, and on Ko-GDPval, an in-house Korean officework-agent benchmark, it is competitive with DeepSeek-V4-Pro (1.6T) at less than a sixth of its size.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑