DAVET:面向扩散视觉-语言模型的降噪感知视觉证据轨迹分配
DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models
浏览论文内容
中文总结 AI 辅助
DAVET 是一种无需训练的框架,通过根据扩散视觉-语言模型的生成状态自适应分配视觉证据,实现平均 1.55 倍加速且仅 1.86% 的平均性能下降,降低了视觉条件成本。
中文摘要 AI 辅助
扩散视觉-语言模型(dVLMs)在每次降噪步骤中都以视觉证据为条件,迭代地对掩码响应进行降噪,这使得视觉条件成为大量重复的推理成本。与自回归解码不同,扩散生成会随着不确定性的演变反复重新处理整个响应。我们的分析表明,视觉证据需求具有很强的步骤依赖性,这促使我们在各个降噪步骤中进行自适应分配。现有的推理加速方法通过解码端策略或通过剪枝与合并进行的视觉 token 压缩来实现,但并未明确将视觉证据视为一种需求随扩散过程演变的资源。因此,我们提出了 Denoising-Aware Visual Evidence Trajectory Allocation(DAVET,降噪感知视觉证据轨迹分配),这是一种无需训练的框架,可根据不断演变的生成状态分配视觉证据。从阶段条件化的证据轨迹开始,所提出的分配策略利用操作需求设置证据储备,其在每个降噪步骤的分配由轨迹风险进行调节。DAVET 通过从单个视觉编码构建的一系列证据视图来实现所得预算,将所需证据的时机与数量和证据视图的构建方式分离开来。在两个代表性 dVLMs(LLaDA-V 和 LaViDa)以及多个视觉理解基准上进行评估,DAVET 实现了平均 1.55 倍的加速,同时平均相对性能下降为 1.86%,表明降噪感知视觉证据分配可降低视觉条件成本,同时在很大程度上保持生成质量。
英文摘要
Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost. Unlike autoregressive decoding, diffusion generation repeatedly revisits the entire response as uncertainty evolves. Our analysis reveals that visual evidence demand is strongly step-dependent, motivating adaptive allocation across denoising steps. Existing inference acceleration methods operate through decoding-side strategies or visual token compression via pruning and merging, but do not explicitly treat visual evidence as a resource whose demand evolves across the diffusion process. Therefore, we present Denoising-Aware Visual Evidence Trajectory Allocation (DAVET), a training-free framework that allocates visual evidence according to the evolving generation state. Starting from a phase-conditioned evidence trajectory, the proposed allocation policy uses operation demand to set an evidence reserve whose allocation at each denoising step is modulated by trajectory risk. DAVET realizes the resulting budgets through a hierarchy of evidence views constructed from a single visual encoding, separating when and how much evidence is needed from how the evidence views are constructed. Evaluated on two representative dVLMs, LLaDA-V and LaViDa, across multiple visual-understanding benchmarks, DAVET achieves an average speedup of 1.55$\times$ with an average relative performance drop of 1.86\%, showing that denoising-aware visual evidence allocation can reduce visual conditioning cost while largely preserving generation quality.