发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过两个时间时钟区分扩散多模态大语言模型中答案稳定与推理展开,发现覆盖率是提示差异的主要分量,并量化了块长度和直接指令对准确率的影响。
AI 中文摘要
在掩码扩散多模态大语言模型中,答案候选可能在推理过程仍在展开时就已经稳定。我们将记录候选的回顾性稳定与令牌承诺区分开来,并考察这两个时钟相对于推理生成的关系。通过分析我们在三个视觉问答基准上的结果,我们发现,在单块、抑制EOS的LaViDa运行中,稳定时推理侧画布仍有89.4%-98.1%未写入。在V*Bench上,将块长度从128减少到8,该比例从89.4%变为1.7%,同时答案覆盖率和符合条件的观察窗口也随之变化。在启用EOS的提示下,直接指令使Nemotron在M3CoT和ScienceQA上的整体准确率分别提高了15.0和19.5个百分点,但使LaViDa/V*Bench的准确率降低了11.0个百分点。对称分解表明,每次变化的较大绝对分量归因于覆盖率而非条件准确率。匹配画布的图像消融实验在答案稳定的同时测量视觉敏感性,将两个时间读数分开。综合这些测量,我们区分了答案稳定、推理展开和视觉敏感性,并确定覆盖率是提示差异的较大分量。
英文摘要
An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On V*Bench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/V*Bench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.
CommentsNeurIPS 2026 Workshop on BeNTo (Beyond Next-Token Prediction - Diffusion & Flow Models for Next-Generation Decoding)