arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Twins:使用焦点损失学习预测统一表示

Twins: Learn to Predict Unified Representations with Focal Loss

Kaixiong Gong, Xin Cai, Bin Lin, Hao Wang, Yunlong Lin, Mingzhe Zheng, Bohao Li, Jian-Wei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Xiangyu Yue

arXiv 2607.22531首次发表:更新:

发表机构

Tencent-Hunyuan(腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究统一多模态模型潜在空间不匹配问题,提出Twins统一连续令牌空间。因联合建模有优化不平衡,采用焦点回归目标解决,在ImageNet上有gFID增益,在多模态理解基准测试中表现好,还提高了重建保真度。

AI 中文摘要

统一多模态模型寻求支持多模态理解和图像生成的共享视觉令牌空间。离散方法通过共享码本统一接口,而连续管道通常依赖两种不同的表示——用于理解的语义特征(如ViT)和用于合成的低级潜在特征(如VAE),导致潜在空间不匹配。我们提出了Twins,一个通过在同一令牌网格上按通道连接ViT和VAE特征形成的统一连续令牌空间,序列长度不变且注意力成本不增加。然而,在扩散变换器中联合建模Twins会出现严重的优化不平衡:模型能很好地拟合ViT组件,但难以匹配VAE潜在分布。我们将这种不平衡追溯到三个异质性来源:频率偏差、内在维度以及条件对齐与条件独立的不确定性。为了解决这个问题,我们采用焦点回归目标进行流匹配,对大误差的VAE维度进行加权,更好地平衡ViT和VAE组件之间的优化。在ImageNet上,与无分类器指导的朴素MSE损失相比,这带来了高达10.57的gFID增益。Twins在多模态理解基准测试中也具有竞争力,并提高了重建保真度,缩小了面向理解和生成的表示之间的差距。

英文摘要

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations -- semantic features (e.g., ViT) for understanding and low-level latents (e.g., VAE) for synthesis -- resulting in mismatched latent spaces. We propose Twins, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase. However, jointly modeling Twins in a Diffusion Transformer exposes a severe optimization imbalance: the model fits the ViT component well but struggles to match the VAE latent distribution. We trace this imbalance to three sources of heterogeneity: frequency bias, intrinsic dimensionality, and condition-aligned vs condition-independent uncertainty. To address it, we adapt a focal regression objective for flow matching that upweights large-error VAE dimensions, better balancing optimization across the ViT and VAE components. On ImageNet, this yields up to 10.57 gFID gain over naive MSE loss without classifier-free guidance. Twins also performs competitively on multimodal understanding benchmarks and improves reconstruction fidelity, narrowing the gap between understanding- and generation-oriented representations.

CommentsICML 2026. Code: https://github.com/Tencent-Hunyuan/Twins

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑