arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文感知集群解码:dMLLMs中的语义锚驱动连贯性

Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs

Yikai Zhao, Qiyan Zhao, Jiaquan Zhang, Xiaofeng Zhang, Xiaosong Yuan, Pengzhou Cheng

arXiv 2608.22367首次发表:更新:

发表机构

Sun Yat-sen University; Shanghai Jiao Tong University; University of Electronic Science and Technology of China; Alibaba Group; Shanghai University(中山大学; 上海交通大学; 电子科技大学; 阿里巴巴集团; 上海大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 dMLLMs 生成长文本时的语义漂移与重复问题,提出上下文感知集群解码方法,结合置信度与邻域邻近度评分,在多模型多基准上实现质量提升并减少幻觉。

AI 中文摘要

扩散多模态大语言模型(dMLLMs)在生成长文本时,常出现语义漂移和重复问题,且输出长度增加时质量通常会下降。我们发现现有解码方法存在两个结构性缺陷是导致这些问题的主要原因:基于置信度的评分忽略了解码邻域的支持,而块划分机制阻碍了对高就绪度语义锚的访问,二者共同导致令牌在其局部上下文未充分建立前就被确定。我们提出了 \textbf{上下文感知集群解码(Context-Aware Cluster Decoding, \textbf{ours})},这是一种无需训练的解码方法,通过将 softmax 置信度与邻域邻近度相乘的复合值对每个掩码位置进行评分,优先选择上下文就绪的令牌而非孤立候选,同时抑制低置信度的位置噪声;该方法采用无块划分机制,使高就绪度的语义锚保持全局可访问。\textbf{ours} 还应用了感知架构的校准,以处理由不同视觉整合策略导致的置信度异质性。在三个 dMLLMs 和四个基准上进行的实验表明,与原始方法相比,该方法实现了持续的质量提升并减少了幻觉,在多个生成长序列的场景中获得了更大的增益,凸显了邻域支持和视觉整合策略对未来 dMLLM 解码方法设计的重要性。我们的代码可在该 https URL 公开获取。

英文摘要

Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies in existing decoding methods as primary drivers of these failures: confidence-based scoring ignores decoded-neighbor support, and block partitioning prevents access to high-readiness semantic anchors, together causing tokens to be committed before their local context is sufficiently established. We propose \ours{} (\textbf{C}ontext-\textbf{A}ware \textbf{C}luster \textbf{D}ecoding), a training-free decoding method that scores each masked position by a multiplicative composite of softmax confidence and neighbor proximity, promoting contextually ready tokens above isolated candidates while suppressing low-confidence positional noise, operating block-free to keep high-readiness anchors globally accessible. \ours{} further applies architecture-aware calibration to handle confidence heterogeneity induced by diverse visual integration strategies. Experiments on three dMLLMs across four benchmarks demonstrate consistent quality gains and hallucination reduction over Original, with larger gains in several longer generation settings, highlighting the importance of neighbor support and visual integration strategy for future dMLLM decoding method design. Our code is openly available at https://github.com/zhaoyk-sysu/CACD-dMLLM.

CommentsThis paper is accepted by EMNLP 2026. 19 pages, 12 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑