梯度冲突能否预测理解-生成权衡?统一多模态模型中冲突度量有效性的受控审计
Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models
浏览论文内容
中文总结 AI 辅助
本研究在受控测试平台GRIDUMM上审计统一多模态模型中梯度冲突度量对理解-生成权衡的预测有效性,发现方向性冲突度量相关性弱且干预无效,表明其诊断有效性需被验证而非假定。
中文摘要 AI 辅助
统一多模态模型(UMMs)的设计日益围绕理解目标与生成目标之间的梯度冲突展开。然而,减少这些度量指标能够改善下游理解-生成权衡这一前提从未被直接检验。我们在一个受控测试平台GRIDUMM中对此进行了审计,该平台镜像了UMM训练的关键结构要素,同时使真实权衡可精确计算。在63种配置和372个测量检查点中,没有任何方向性冲突度量在训练期间测量的冲突与最终权衡之间达到绝对Spearman相关系数0.3且置信区间排除零。一个单调抑制冲突的剂量-反应干预使权衡保持平坦,从而将相关性与因果性分离开来。范数比率是一个生成失败检测器,在掌握生成的配置中变为无效。功能性干扰度量优于方向性冲突度量,而训练损失则强有力地追踪权衡。我们的结果并未表明冲突无用;而是表明其作为诊断目标的有效性必须被确立,而非被假定,我们发布了审计协议作为可复用的标准。
英文摘要
Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly. We audit it in a controlled testbed, GRIDUMM, which mirrors key structural ingredients of UMM training while making the ground-truth trade-off exactly computable. Across 63 configurations and 372 measured checkpoints, no directional conflict metric reaches an absolute Spearman correlation of 0.3 with a confidence interval excluding zero for conflict measured during training against the eventual trade-off. A dose-response intervention that monotonically suppresses conflict leaves the trade-off flat, separating correlation from causation. The norm ratio is a generation-failure detector and becomes null among configurations that master generation. Functional interference measures outperform directional conflict metrics, while training loss tracks the trade-off strongly. Our results do not show that conflict is useless; they show that its validity as a diagnostic target must be established, not assumed, and we release the audit protocol as a reusable standard.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
- Aimakj
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。