arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信息引导的前沿解码:扩散多模态语言模型(dMLLMs)中上下文效用驱动的提交

Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs

Xingyou Fang, Jingxing Zhong, Xiaosong Yuan, Xiaofeng Zhang

arXiv 2608.26641首次发表:更新:

发表机构

Fuzhou University; Jilin University; Shanghai Jiao Tong University(福州大学; 吉林大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散多模态语言模型解码时结构令牌过早提交削弱上下文传播的问题,提出无需训练的信息引导前沿解码策略,通过多维度排序优化候选选择,在多类基准上实现了优于现有方法的性能。

AI 中文摘要

扩散多模态语言模型(dMLLMs)的解码质量高度依赖于掩码令牌的提交顺序。现有的基于置信度的策略优先选择局部易处理的令牌,但置信度并不一定反映上下文有用性。因此,标点等结构上易处理的令牌可能会在信息丰富的语义锚点之前被提交,削弱上下文传播并增加错误累积。我们提出信息引导的前沿解码(IGFD),这是一种无需训练的解码策略,通过令牌置信度、邻域不确定性和结构提交风险对候选进行排序。IGFD鼓励早期提交可靠的语义锚点,同时延迟脆弱的结构令牌,以改善解码过程中的上下文支持。动态候选前沿进一步将令牌选择限制在相同解码预算下可局部扩展的区域。该方法无需额外训练、辅助模型或额外前向传播。在多模态理解、推理、 grounding(接地)和幻觉基准上的实验表明,在相同解码预算下,IGFD在大多数基准和扩散MLLM主干上始终优于现有解码策略。

英文摘要

Decoding quality in diffusion multimodal language models (dMLLMs) depends heavily on the order in which masked tokens are committed. Existing confidence-based strategies prioritize locally easy tokens, but confidence does not necessarily reflect contextual usefulness. As a result, structurally easy tokens such as punctuation may be committed before informative semantic anchors, weakening context propagation and increasing error accumulation. We propose Information-Guided Frontier Decoding (IGFD), a training-free decoding strategy that ranks candidates using token confidence, neighborhood uncertainty, and structural commitment risk. IGFD encourages early commitment of reliable semantic anchors while delaying fragile structural tokens, improving contextual support during decoding. A dynamic candidate frontier further constrains token selection to locally expandable regions under the same decoding budget. The method requires no additional training, auxiliary models, or extra forward passes. Experiments across multimodal understanding, reasoning, grounding, and hallucination benchmarks show that IGFD consistently outperforms existing decoding strategies across the majority of benchmarks and diffusion MLLM backbones under identical decoding budgets.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑