WeAgent-MMGenEdit:多模态智能体图像生成与编辑的全栈方案
WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing
- Weixin AI, Tencent(腾讯微信人工智能)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
WeAgent-MMGenEdit针对多模态智能体图像生成编辑的局限,提出含WeAgent-Harness等的全栈方案,构建基准与训练方法,使30B总规模策略性能接近1T参数智能体
AI中文摘要:
图像生成与编辑模型发展迅速,但在提示词需要外部世界知识时仍不可靠。有限且长尾的参数知识使得直接生成或推理后生成的方法无法恢复所需事实与视觉外观。现有智能体生成与编辑方法通过检索工具缓解此局限,但仍受视觉验证不足、策略模型过载、检索到的文本与视觉证据集成薄弱的限制。为解决这些局限,本文提出WeAgent-MMGenEdit,这一全栈方案包含多模态操控器、可扩展数据构建流水线、综合基准及针对智能体策略与图像后端的后训练方法。首先介绍WeAgent-Harness,这一具有持久证据管理及专用验证与集成工具的多模态运行时,将检索到的多模态证据组织为密集载体。在此基础上,开发用于提示词合成与智能体轨迹收集的可扩展流水线,产出23K条监督轨迹及14.7K个带三层可验证清单的强化学习任务。进一步推出WeBench-MMGenEdit,这一覆盖知识密集型图像生成与多图像编辑的双语基准。最后,基于SFT与RL的双边后训练方案优化智能体策略与图像后端。总体而言,WeAgent-MMGenEdit使总规模30B、激活规模3B的策略模型,能超越同规模策略模型,接近1T参数智能体的性能。
英文摘要:
Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approaches from recovering the required facts and visual appearances. Existing agentic generation and editing methods mitigate this limitation with retrieval tools, yet remain constrained by insufficient visual verification, overloaded policy models, and weak integration of retrieved textual and visual evidence. To address these limitations, we present WeAgent-MMGenEdit, a full-stack recipe including a multimodal harness, a scalable data construction pipeline, a comprehensive benchmark, and post-training methods for the agent policy and image backend. We first introduce WeAgent-Harness, a multimodal runtime with persistent evidence management and dedicated verification and integration tools that organize retrieved multimodal evidence into a dense carrier. Upon this, we develop a scalable pipeline for prompt synthesis and agentic trajectory collection, yielding 23K supervised trajectories and 14.7K RL tasks with three-layer verifiable checklists. We further introduce WeBench-MMGenEdit, a bilingual benchmark covering both knowledge-intensive image generation and multi-image editing. Finally, a two-sided post-training recipe based on SFT and RL improves the agent policy and image backend. Together, WeAgent-MMGenEdit enables a 30B-total/3B-active policy to outperform similarly sized policy models and approach the performance of a 1T-parameter agent.