更快但不同:加速多模态扩散语言模型中的内容漂移诊断与控制
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对加速多模态扩散语言模型的内容漂移问题,通过实验诊断出漂移源于过期状态,提出缩短KV缓存刷新区间的控制方法,在1.3倍加速时实现近乎完全一致,且未发现事实错误差异。
中文摘要 AI 辅助
无需训练的加速方法使基于扩散的多模态大语言模型(dMLLMs)更易部署,但可能悄悄改变生成内容。我们针对300张真实图像研究该推理阶段一致性问题,对比Fast-dLLM输出与同一模型未加速输出。在长文本设定下,轻度并行导致每步提交token数为1.05-1.25,置信度阈值调整会改变解码行为但不影响基线一致性;状态刷新消融实验和图像交换干预则发现,过期的视觉和生成文本状态是漂移的原因。对于测试的Fast-dLLM实现,缩短KV缓存刷新区间会产生单调的速度-一致性边界,在1.3倍加速时达到近乎完全一致。该初始诊断也出现在dLLM-Cache和LaViDa中,不过dLLM-Cache仅在收紧两个缓存后恢复一致性,这会消除其速度优势。独立的提示词和图像可复现阈值不敏感性和刷新恢复特性。针对性审计发现,50对低一致性样本中有一半存在真实内容替换。在另一项双标注员盲评中,加速组与基线组的事实错误差值为0.00(95%置信区间[-0.17,+0.17]);该样本未检测到差异,但不代表事实等价。最后,测试的自适应或平滑刷新变体在相同计算量下均未优于固定区间。我们的贡献是一套配对诊断方法和实现范围内的一致性控制方案,而非准确率或安全保证。
英文摘要
Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing Fast-dLLM outputs with the same model's unaccelerated outputs. Across the mild parallelism induced in our long-form setting (1.05--1.25 committed tokens per step), confidence-threshold tuning changes decoding behavior but not baseline agreement. State-refresh ablations and an image-swap intervention instead identify stale visual and generated-text states as contributors to drift. For the tested Fast-dLLM implementation, shortening the KV-cache refresh interval yields a monotonic speed--agreement frontier and near-exact agreement at a measured 1.3x speedup. The initial diagnosis also appears with dLLM-Cache and LaViDa, although dLLM-Cache recovers agreement only after both caches are tightened, which removes its speed advantage. Independent prompts and images reproduce the threshold-insensitivity and refresh recovery. A targeted audit finds genuine content substitution in half of 50 low-agreement pairs. In a separate blinded two-annotator evaluation, the pooled accelerated-minus-baseline factual-error difference is 0.00 (95% CI [-0.17,+0.17]); this sample detects no difference but does not establish factual equivalence. Finally, none of the tested adaptive or smoothed-refresh variants beats the fixed interval at matched compute. Our contribution is a paired diagnostic and an implementation-scoped consistency control, not an accuracy or safety guarantee.
发表机构
- School of Mathematics and Statistics, Beijing Institute of Technology(北京理工大学数学与统计学院)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。