相互对抗自训练与演化数据用于统一多模态模型
Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models
- Georgia Institute of Technology(佐治亚理工学院)
- Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MATE框架,通过生成与理解分支相互对抗并演化挑战数据,强化统一多模态模型,提升生成与理解基准性能。
AI中文摘要:
统一多模态模型(UMMs)在共享骨干网络中结合了图像生成和视觉理解。由于生成和理解是逆任务,近期研究通过让两个分支相互协作监督来对UMMs进行自训练。我们提出了MATE(相互对抗自训练与演化数据),一种基于强化学习的后训练框架,其中两个分支相互挑战,并且随着模型训练,挑战不断演化。MATE让生成和理解轮流担任挑战者和求解者。给定一张图像,理解分支提出多个候选描述,生成分支必须将这些描述还原为相似的图像,反之亦然。候选描述会根据其来源的图像或提示进行一致性筛选,并且求解器在其处理最差的候选上进行训练。因此,对抗性来自模型自身的输出,无需训练单独的对抗器。此外,在下一个epoch中,击败一个分支的候选成为对另一个分支的下一个挑战的来源,这使得挑战随模型演化,并将训练转化为数据空间中的自我博弈。在Janus-Pro-1B上,MATE将GenEval提升2.4个百分点,DPG-Bench提升1.7个百分点,九个理解基准的平均分提升0.7个百分点,同时增强了重复图像-文本循环中的一致性。
英文摘要:
Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.