arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11977cs.AI

Occamy-1.0:面向协同工作的开放帕累托前沿35B智能体

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

  • Accio Team(Accio 团队)

机构由 AI 辅助整理,请以论文原文为准。

Wenhui Chen, Shiwen Cheng, Hao Dong, Chenda Duan, Ruixiang Feng, Zhong Guan, Boqiang Guo, Xueyuan Han, Haojie Hao, Liangmeng Huang, Zhelong Huang, Xinke Kong, H… 展开作者

Wenhui Chen, Shiwen Cheng, Hao Dong, Chenda Duan, Ruixiang Feng, Zhong Guan, Boqiang Guo, Xueyuan Han, Haojie Hao, Liangmeng Huang, Zhelong Huang, Xinke Kong, Hongyu Li, Jiazheng Li, Junbo Li, Qingchuan Li, Yukun Lian, Chang Liu, Tianyu Liu, Zicheng Liu, Shuyi Ouyang, Yijun Pan, Kunyu Shi, Xiaojun Tang, Bingquan Wang, Kesu Wang, Yuchen Wang, Sibo Wei, Sicong Xie, Xiaoying Xing, Yi Xu, Zhijun Xu, Hongwei Xue, Qingcheng Zeng, Di Zhang, Guannan Zhang, Haochen Zhang, Tianlong Zhang, Tianyu Zhao, Tianyu Zhao, Yanjun Zheng, Jialong Zhu, Zijian Zou

AI总结:

Occamy-1.0通过对Qwen3.6-35B-A3B进一步后训练,构建执行数据与多框架轨迹,实现成本高效的协同工作智能体,在基准上达到帕累托前沿低成本拐点并保持广泛能力。

AI中文摘要:

协同工作智能体执行复杂的工作流程,这些流程在多次模型调用中结合了信息收集、工具使用、编码和文件操作。由于成本和延迟在整个回合中累积,其实际价值不仅取决于峰值能力,还取决于该能力被交付的效率。然而,日常工作中的许多步骤强调状态跟踪、协调、恢复和跟进,而非前沿规模的推理。我们提出了Occamy-1.0,一种成本高效的协同工作模型,通过对后训练的Qwen3.6-35B-A3B检查点进行进一步训练而获得。我们构建了基于执行的数据和环境,捕获了跨多个框架的可重放的长视野轨迹,并使用分阶段的后训练来发展和巩固互补的执行能力。在广泛的协同工作基准测试套件中,Occamy-1.0始终是同等规模模型中最强的之一,并在多项任务上与规模大得多的前沿系统保持竞争力。在我们声明的评估和定价协议下,其在四个代表性基准上的总体性能使其处于观察到的成本-性能帕累托前沿的低成本拐点。支持性的工具调用、编码和指令遵循评估进一步表明,这种专业化保留了广泛的智能体能力。我们发布了模型权重和一部分训练数据,以支持对实用协同工作智能体和智能体后训练的研究。

英文摘要:

Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.

↑