arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29505cs.LG

谱引导扩散:通过静态谱层调度加速推理

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

  • Iowa State University(爱荷华州立大学)
  • BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma

AI总结:

本文提出谱引导扩散方法,利用谱集中比与Frobenius范数离线调度冻结网络层,在不需路由器或输入搜索下加速扩散推理,在多个模型上实现2.8-3.0倍墙钟加速并保持生成质量。

AI中文摘要:

扩散推理会重复评估同一个大型网络。我们探究是否仅凭预训练权重就能识别出在整个轨迹中无需重新计算的残差分支。我们提出的谱集中比(SCR)衡量前导奇异值与尾部奇异值能量之比。结合Frobenius范数,它产生一个离线敏感性代理指标,并为每个被调度的单元确定一个确定性的生命周期。一个被冻结的单元会重用其缓存的残差分支更新,而当前的残差流和所有外部条件继续传播。该方法不需要路由器、校准提示或依赖输入的搜索。在匹配的层-步预算下,SCR/Frobenius在LLaDA-8B、DiT-XL/2、U-ViT-L和SDXL上比随机、深度、范数、稳定秩和Frobenius-稳定秩调度更好地保持质量。更广泛的LLaDA测试涵盖检索、推理、代码、摘要和开放式生成;匹配视野的控制在低至十步去噪时仍保持排名。完整的捕获图系统相比急切推理实现了2.8倍至3.0倍的墙钟加速。这是一个系统层面的数字:在LLaDA上,填充图执行已经给出2.7倍加速,而消除非活动分支工作将其提升至3.0倍。扰动分析在明确的局部假设下为预归一化注意力机制和MLP组件提供了动机;在AdaLN、U形、卷积和交叉注意力块上的结果是经验迁移,而非认证保证。

英文摘要:

Diffusion inference repeatedly evaluates the same large network. We ask whether pretrained weights alone can identify residual branches that need not be recomputed throughout the trajectory. Our \textbf{Spectral Concentration Ratio (SCR)} measures leading-versus-tail singular-value energy. Combined with Frobenius magnitude, it yields an offline sensitivity proxy and a deterministic lifetime for each scheduled unit. A frozen unit reuses its cached residual-branch update while the current residual stream and all external conditioning continue to propagate. The method needs no router, calibration prompts, or input-dependent search. At matched layer-step budgets, SCR/Frobenius preserves quality better than random, depth, norm, stable-rank, and Frobenius--stable-rank schedules on LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL. Broader LLaDA tests cover retrieval, reasoning, code, summarization, and open-ended generation; matched-horizon controls retain the ranking down to ten denoising steps. The complete captured-graph system reaches $2.8\times$--$3.0\times$ wall-clock speedup over eager inference. This is a systems-level number: on LLaDA, padded graph execution already gives $2.7\times$, while eliminating inactive branch work raises it to $3.0\times$. The perturbation analysis motivates pre-norm attention and MLP components under explicit local assumptions; results on AdaLN, U-shaped, convolutional, and cross-attention blocks are empirical transfer, not certified guarantees.

↑