arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19878cs.LGcs.CL

Uni-LaDiR:潜在扩散统一多模态推理

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Jiatao Gu, Yian Ma, Lianhui Qin

首次发表
浏览论文内容

中文总结 AI 辅助

Uni-LaDiR提出统一潜在扩散推理框架,将多模态思维映射至共享空间,在11个VLM基准和2个VLA套件上分别相对提升7.3%和6.1%。

中文摘要 AI 辅助

多模态推理要求模型在整个推理过程中利用来自多种模态的信息。然而,现有方法通常将特定模态的思维标记拼接在单一序列中,使得模型在跨模态推理时不得不自行弥合表征差异。我们提出了Uni-LaDiR(统一潜在扩散推理器),一个将这些思维引入共享潜在空间进行推理的框架。一个统一的编码器将来自不同模态的教师推理步骤映射为共享思维标记,并训练其保留后续推理步骤及最终答案或动作所需的信息。由于相同的上下文可以支持多个有效的后续步骤,我们使用扩散模型从输入和前面的块中预测下一个思维标记块。通过共享模型权重联合训练编码器和扩散推理器,鼓励思维标记既对任务有用,又能从可用上下文中预测。在推理时,模型无需教师观测即可生成这些标记。在十一个视觉语言模型(VLM)基准和两个视觉-语言-动作(VLA)套件上,Uni-LaDiR相对于最强评估基线在视觉推理任务上取得了7.3%的相对提升,在机器人操作任务上取得了6.1%的相对提升。

英文摘要

Multimodal models increasingly think with different modalities such as images, 3D point clouds, and robot states, not just text. Yet each modality is still encoded into its own representation space, creating a modality-switching gap whenever reasoning moves from one modality to another. In this paper, we introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that unifies different modalities into a shared latent space for multimodal reasoning. A unified encoder maps teacher reasoning steps from different modalities into latent thought tokens in a shared space, trained to extract the information needed for later reasoning steps and the final output. A diffusion reasoner, trained jointly with the encoder, generates these tokens at inference without teacher reasoning steps. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest baselines of 7.3% on four mathematical and logical VLM benchmarks and 6.1% on RLBench manipulation tasks. Controlled comparisons show increasing gains as more teacher modalities are unified. These results suggest that unification improves multimodal reasoning by weaving it into a single thread, where the model predicts successive thoughts in a common representation space.

发表机构

  • UC San Diego(加州大学圣迭戈分校)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

↑