Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
理解与生成相冲突吗?统一多模态模型DPO的诊断研究
专题命中 偏好对齐 :DPO(title,title_cn);alignment(abstract);分类 cs.AI、cs.LG
AI总结 通过系统实验发现,在统一多模态模型上应用DPO时,生成质量难以对齐,主要原因是理解和生成梯度近乎正交且存在11-14倍的幅度不平衡,源于VQ token数量不对称。
Comments Experiments are inconclusive: The claim that architectures such as Chameleon or Emu would exhibit stronger gradient conflict is not supported by experiments or analysis, and all experiments are conducted on Janus-Pro without evaluation on other unified multimodal architectures