arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从零步洞察终点:通过MLP稀疏感知截断加速扩散多模态大语言模型

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation

Qicheng Zhao, Qi Sun, Zheyu Yan

arXiv 2607.14557首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散多模态大语言模型推理效率受固定长度生成约束影响的问题,提出Seer框架,利用MLP激活稀疏性变化检测有效语义边界,进行冗余截断及采用混合执行策略,实验表明其可加速吞吐量并维持性能甚至提升复杂视觉任务准确性。

AI 中文摘要

扩散多模态大语言模型(DMLLMs)在多模态推理中非常有效,但其推理效率受到固定长度生成约束的显著阻碍。由于实际输出长度未知,输出序列被填充到预定义的最大长度,导致对不必要的[EOS]令牌进行大量冗余计算。在这项工作中,我们发现DMLLMs在第一个去噪步骤通过MLP激活稀疏性的明显变化隐含地揭示其有效语义边界。利用这一观察结果,我们提出了Seer,一个无需训练的框架,它使用基于信噪比(SNR)的标准检测这个边界,并对所有后续计算的冗余后缀进行一次性截断。为了在批量服务期间保持这些理论收益,Seer采用了一种混合执行策略,在无缝适应动态序列长度的同时最大化吞吐量。实验结果表明,Seer有效地消除了填充浪费,吞吐量提高了约31倍。在9个基准测试中,Seer稳健地保持整体性能,甚至通过减轻噪声泄漏提高了复杂视觉任务的准确性(例如,DocVQA分数从63 .52提高到63 .66),为DMLLM加速提供了一种高效、即插即用的解决方案。

英文摘要

Diffusion Multimodal Large Language Models (DMLLMs) are highly effective for multimodal reasoning, yet their inference efficiency is significantly hindered by fixed-length generation constraints. Since the actual output length is unknown, output sequences are padded to a predefined maximum length, resulting in substantial redundant computation over unnecessary [EOS] tokens. In this work, we discover that DMLLMs implicitly reveal their valid semantic boundary at the very first denoising step through a distinct shift in MLP activation sparsity. Leveraging this observation, we propose Seer, a training-free framework that detects this boundary using a Signal-to-Noise Ratio (SNR)-based criterion and performs one-shot truncation of the redundant suffix for all subsequent computations. To preserve these theoretical gains during batched serving, Seer incorporates a hybrid execution strategy that maximizes throughput while seamlessly accommodating dynamic sequence lengths. Experimental results demonstrate that Seer effectively eliminates padding waste, accelerating throughput by up to $\sim$31$\times$. Across 9 benchmarks, Seer robustly maintains overall performance and even improves accuracy on complex visual tasks by mitigating noise leakage (e.g., DocVQA score increases from 63.52 to 63.66), offering a highly efficient, plug-and-play solution for DMLLM acceleration.

CommentsAccepted to ACM Multimedia (ACM MM) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑