arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38721cs.AIcs.CV

UniEvo-VL:一种用于多模态模型自我提升的在线策略自蒸馏训练方案

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

  • Stanford University(斯坦福大学)
  • Johns Hopkins University(约翰霍普金斯大学)
  • University of Toronto(多伦多大学)
  • University of Oxford(牛津大学)
  • UC, Riverside(加州大学河滨分校)
  • MatrAIx
  • University of Notre Dame(圣母大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • UNC–Chapel Hill(北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng X… 展开作者

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi

AI总结:

提出UniEvo-VL,一种在线策略自蒸馏训练方案,让多模态模型同时作为教师和学生,利用自我批评反馈提升图像生成能力,实验显示在GenEval等基准上显著性能提升。

AI中文摘要:

现代多模态模型将生成与理解集成于一个统一系统中,使其能够提供并学习自身的反馈。受这种统一能力的启发,我们提出了UniEvo-VL,一个自进化框架,使多模态模型在测试时计算过程中能够从这种建设性的自我纠正反馈中学习。我们不再依赖一个单独的、通常更大的教师模型,而是利用其自我批评作为特权信息,并让同一个多模态模型在不同上下文中同时扮演教师和学生角色。学生仅看到原始问题,而教师则基于特权批评进行条件化。随后,训练最小化学生自身采样轨迹上其去噪扩散分布之间的逐状态差异。实验表明,UniEvo-VL提升了多模态模型的图像生成能力,同时保持其对额外反思信息的敏感性。具体而言,我们基于开源Qwen-image-2512构建,观察到在GenEval上性能从0.747显著提升至0.808,在GenEval2 Soft-TIFA上从32.97提升至35.53。此外,尝试使用更强大的外部批评者(如GPT5.6-Luna)表明,具有强大评判能力的多模态模型可以预期更高的自进化上限。最后但同样重要的是,混合文本渲染结果表明,我们的自我改进可能在不同任务间并不均匀。我们的研究旨在为当前热门的递归自我改进研究路线提供启示,以在无需外部监督或指导的情况下提升多模态模型的使用体验。

英文摘要:

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.

↑