AI 中文总结
本文介绍为长期代理任务构建的太阳能开放2语言模型。通过混合注意力堆栈等技术扩大规模,利用更强起点和高价值数据实现高效训练,经多教师策略蒸馏构建代理技能,在多项基准测试中表现出色,展示了其在语言模型领域的优势。
AI 中文摘要
我们展示了太阳能开放2,这是一个为长期代理任务构建的250B - A15B专家混合语言模型,从太阳能开放1(太阳能开放100B)扩展而来。为在单个上下文中保存整个代理轨迹,太阳能开放2通过混合注意力堆栈达到1M令牌窗口,该堆栈在每三个线性注意力层中交错一个softmax层,不使用位置编码和扩展到负特征值的门控增量规则。为在固定计算预算下进行这种规模的训练,我们通过更强的起点和更高价值的数据使训练高效。起点方面,从太阳能开放1初始化,转移架构变化后幸存的56.9亿参数共享框架并通过全预训练学习其他内容。数据方面,按每个令牌的价值精心挑选,通过质量和稀有度感知的数据策划和混合比例优化将20T池细化为10T混合物,在相同令牌预算下优于太阳能开放1的方法。为构建其代理技能,我们在专门场景中训练12个领域专家,然后通过多教师策略蒸馏(MOPD)将它们整合到单个模型中。在英语基准测试中,太阳能开放2在MMLU - Pro、LiveCodeBench和APEX - Agents代理套件上领先,在其他地方与最强模型(DeepSeek - V4 - Flash和MiMo - V2.5)竞争。在韩语基准测试中,太阳能开放2在所有比较模型中平均得分最高,包括快速层级的封闭API,在内部韩语办公代理基准测试Ko - GDPval上,它以不到DeepSeek - V4 - Pro(1.6T)六分之一的规模与之竞争。
英文摘要
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, we make training efficient in two ways: a stronger starting point, and higher-value data. For the starting point, we initialize Solar Open 2 from Solar Open 1, transferring the 5.69B-parameter shared skeleton that survives the architectural change and learning everything else through full pre-training. For the data, we curate for value per token: quality- and rarity-aware data curation and mixture-ratio optimization refine a 20T pool into a 10T mixture that, at equal token budget, outperforms the Solar Open 1 recipe. To build its agent skills, we train twelve domain specialists across purpose-built scenarios, then consolidate them into a single model by Multi-teacher On-Policy Distillation (MOPD). Against comparably sized open-weight models on English benchmarks, Solar Open 2 leads on MMLU-Pro, LiveCodeBench, and the APEX-Agents agentic suite, and stays competitive with the strongest (DeepSeek-V4-Flash and MiMo-V2.5) elsewhere. On Korean benchmarks, Solar Open 2 records the highest average of any model compared, including fast-tier closed APIs, and on Ko-GDPval, an in-house Korean officework-agent benchmark, it is competitive with DeepSeek-V4-Pro (1.6T) at less than a sixth of its size.