arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01917cs.CVcs.AIcs.CL

MoLE: 用于互补视觉推理的潜在专家混合模型

MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin

首次发表
浏览论文内容

中文总结 AI 辅助

MoLE通过潜在专家混合框架,隔离并聚合互补视觉信息,在五个视觉推理基准上平均得分78.6,优于现有方法,证明专门化潜在计算优于增加潜在标记数量。

中文摘要 AI 辅助

潜在视觉推理为视觉-语言模型配备了连续的中间状态,这些状态能够处理视觉证据,而无需显式的文本推理轨迹或重复的图像操作。然而,现有方法通常允许多个潜在标记通过共享的值投影访问相同的视觉证据,这为它们提取互补的视觉信息提供了机制;因此,简单地增加潜在预算可能会导致冗余的潜在表示。我们认为,有效的潜在推理应鼓励不同的潜在标记提取互补的视觉信息,从而充当专门的视觉专家。基于这一见解,我们提出了MoLE,一个潜在专家混合框架,它控制每个潜在视觉专家观察哪些视觉证据以及如何转换该证据。MoLE在证据提取期间隔离潜在视觉专家,并使用专门的潜在摘要专家来聚合潜在视觉专家的互补表示。两阶段训练流程首先强制视觉证据通过这条潜在路径,然后恢复直接视觉访问,既不需要预定义的专家角色,也不需要中间的视觉目标。在五个视觉推理基准上,MoLE取得了78.6的平均得分,比数据匹配的监督微调高出4.9,比在相同潜在预算下评估的最强潜在视觉推理基线高出3.6。表示分析显示,潜在状态相似性更低,视觉注意力更加多样化,而屏蔽潜在路径会使平均性能降低9.2。这些结果表明,专门化潜在计算比仅仅增加潜在标记的数量更有效。

英文摘要

Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.

发表机构

  • Tongji University(同济大学)
  • Hong Kong Polytechnic University(香港理工大学)
  • Alibaba Group(阿里巴巴集团)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑