Macaron-V1:面向具备自我改进与LoRA混合能力的开放持续学习
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
浏览论文内容
中文总结 AI 辅助
本研究提出Macaron-V1开放智能体模型家族,结合MoL架构与递归自我改进机制,经多基准评估验证其有效性,为开放持续学习提供新方案。
中文摘要 AI 辅助
Macaron-V1是一款面向体验智能的开放智能体模型家族,旨在从真实环境的经验中学习,并在部署后持续学习。它围绕两个系统目标构建:一是通过对带版本的模型-工具对进行递归改进来实现适应,即对某一配置的经验在外部契约下进行评估,用于构建其后续版本;二是通过Mixture-of-LoRA(MoL)架构实现协作,该架构冻结基础模型,组合专用LoRA适配器,并在每一轮用户交互中选择一个LoRA。旗舰版Macaron-V1-Venti结合了7440亿参数的GLM-5.2基础模型与四个针对对话、智能体、编码和GenUI的LoRA;基于Qwen3.6的Macaron-V1-Tall(500亿参数)采用相同设计用于本地部署。本报告将Macaron-V1呈现为跨架构、算法与基础设施的协同设计系统。MoL架构通过可扩展的LoRA专用模块支持持续学习,算法结合了模型-工具协同设计与递归自我改进循环,包括组件原生的GenUI工具UI4A、有状态动作基底、带版本的HCP契约以及智能体强化学习框架MindForge。支撑基础设施包括后训练平台MinT、长上下文强化学习方法LongStraw,以及针对稀疏MoE和DSA基础模型的稳定性技术。我们在个人智能、GenUI和通用能力基准上对Macaron-V1与前沿基线进行评估,结果验证了当前系统的有效性,而持续学习与集体智能的复合增益仍是待解决的问题。
英文摘要
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti (748B) combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-35B-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned Harness Context Protocol contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.
发表机构
- Mind Lab(心智实验室)
机构由 AI 辅助整理,请以论文原文为准。