发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MintAct是一个统一视觉语言模型系列,通过可扩展环境和异步强化学习框架,在2B至8B规模下统一UI定位、多步导航和视觉工具使用,性能达到领域专用模型水平,并在OSWorld-Verified上取得48.9的最优结果。
AI 中文摘要
我们提出了MintAct,一个视觉语言模型系列,它将UI定位、跨移动端、桌面端和网页的多步导航以及视觉工具使用统一起来,并在2B、4B和8B规模上进行训练。通过对环境、数据和训练方案的精心设计,MintAct模型在所有上述能力上均达到了各领域专用模型的性能水平。为实现这一目标,我们开发了可扩展的环境和强化学习(RL)基础设施。在环境方面,我们在异构的领域专用后端上托管数百个并发实例,同时服务于轨迹数据收集和在线强化学习。为了实现高效且可扩展的强化学习训练,一个异步框架明确控制跨领域训练分布,并在嘈杂的环境反馈和离策略漂移下保持稳定。实验结果表明,在可比模型规模下,MintAct在广泛基准测试中达到了最先进的性能(在OSWorld-Verified上为48.9)。
英文摘要
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.