arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向扩散多模态大语言模型的视觉信息引导并行解码

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim

arXiv 2608.26580首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散多模态大语言模型解码时未充分利用输入图像信息的问题,提出视觉信息引导采样器 VIG-Sampler,在 7 个基准及 3 个开源模型上验证其性能优于 Info-Gain 采样器。

AI 中文摘要

扩散多模态大语言模型(dMLLMs)是近期出现的一种多模态生成解码范式,其从完全掩码的序列出发,每一步通过取消部分剩余掩码位置的掩码来逐步解码序列。由于所选 token 会作为后续步骤的预测上下文,因此决定解码哪些 token 对最终输出的质量至关重要。最常见的策略是基于确定性度量对 token 进行优先级排序,该度量往往偏向训练数据中频繁出现的 token。近期的方法则根据 token 对后续预测的影响对其排序,但未明确考虑输入图像。我们提出视觉信息引导采样器(VIG-Sampler),其基于 token 对图像 token 的注意力来对 token 进行优先级排序。我们进一步施加约束,惩罚那些图像注意力分布与先前所选 token 相似的候选 token,从而增加解码子集的信息增益。在 7 个图像描述和视觉问答(VQA)基准、3 个开源 dMLLMs 上进行的大量实验表明,VIG-Sampler 具有有效性:在图像描述基准上,其平均比 Info-Gain 采样器高出 19.3 个 CIDEr 分数,且在 COCO 图像描述任务上超越该采样器,同时仅使用其一半的解码步骤。

英文摘要

Diffusion multimodal large language models (dMLLMs) have recently emerged as a new decoding paradigm for multimodal generation. Starting from a fully masked sequence, dMLLMs progressively decode the sequence by unmasking a subset of the remaining masked positions at each step. Since the selected tokens serve as the prediction context for subsequent steps, deciding which tokens to decode is crucial to the quality of the final output. The most common strategy prioritizes tokens based on a certainty measure that tends to favor tokens frequently observed in the training data. Recent approaches instead order tokens according to their influence on subsequent predictions, but do not explicitly account for the input image. We propose the Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens. We further impose a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded subset. Extensive experiments on 7 captioning and VQA benchmarks with 3 open-source dMLLMs demonstrate the effectiveness of VIG-Sampler, which outperforms the Info-Gain Sampler by an average of 19.3 CIDEr points across the captioning benchmarks and surpasses it on COCO Caption while using only half as many decoding steps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑