arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen-Planner-Agent:面向真实世界移动端规划智能体的闭环AI-for-AI框架

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Tingyu Qu, Weigao Sun, Yuecheng Liu, Yucheng Zhao, Yi Zhu, Yifeng Ding, Qiyi Wang, Sihan Cao, Pengkun Jiao, Hanlei Xie, Xiongwei Wu, Qichao Wang, Haodong Zhang, Jiajun Liu, Yuhao Wang, Yuqing Xie, Junpeng Zhao, Long Chen, Ming Ma, Sihan Yang, Ziwang Zhao, Yanhao Jia, Liangquan Gong, Feida Zhu, Yiran Zhong, Steven Hoi

arXiv 2609.29892首次发表:更新:

发表机构

Alibaba Group; Alibaba Token Hub(阿里巴巴集团; 阿里巴巴Token Hub)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Qwen-Planner-Agent,一种闭环AI-for-AI框架,通过数据飞轮、混合强化学习和执行证据驱动循环,在MobilePA-Bench上取得最佳性能,并提升移动端规划智能体的工具使用与协调能力。

AI 中文摘要

大语言模型的快速发展正将AI从被动的内容生成扩展到工程和科学发现的主动工作流中。这一转变引发了一个引人深思的问题:AI能否既是开发的对象,又是构建下一代AI系统的积极参与者?我们通过在闭环AI-for-AI框架中构建Qwen-Planner-Agent来探索这一问题,以实现可扩展的开发与迭代改进。移动端规划为这一方法提供了严苛的测试:复杂、长时程的任务挑战了智能体的可靠性,而昂贵的真实设备交互限制了开发的可扩展性。该框架通过共享的动作-反馈-验证契约连接数据生产、模型训练和部署。(i)AI for Data构建了一个人工把关的智能体数据飞轮,其中专门的智能体构建任务、收集交互轨迹、整理和平衡训练数据,并利用训练反馈指导后续数据生成。(ii)AI for Training将监督式规划冷启动与混合环境的在线智能体强化学习相结合,我们引入了能力感知的奖励与优势工程(CARE)来降低推理和工具使用成本,同时保持任务性能。(iii)AI通过执行证据驱动的循环驱动模型-工具协同进化,该循环在运行时编排记忆、技能和工具,并将结构化的动作反馈和保留的失败轨迹反馈到协调的模型和工具适配中。Qwen-Planner-Agent在MobilePA-Bench上所有评估模型和系统中取得了最佳整体性能,在工具使用、记忆、技能和子智能体协调方面均优于其基础模型。对我们模型的进一步评估显示,在非移动智能体基准上也有改进,同时基本保持通用能力。

英文摘要

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.

Commentshttps://tongyi-mai.github.io/Qwen-Planner-Agent/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑