arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33723cs.CV

GeoShrink:用两行代码加速扩散Transformer

GeoShrink: Accelerating Diffusion Transformers with Two Lines of Code

Haosen Li, Wenshuo Chen, Shaofeng Liang, Lei Wang, Bowen Tian, Yutao Yue

首次发表
浏览论文内容

中文总结 AI 辅助

GeoShrink是一种无需训练的扩散Transformer加速方法,通过几何锚点预测跳过模型评估,在约5倍加速下显著提升图像、视频、运动、音频和3D生成的保真度。

中文摘要 AI 辅助

扩散Transformer在采样轨迹上通过重复的模型评估产生了大量的推理成本。我们引入了GeoShrink,一种无需训练的加速方法,它在保留原始求解器网格的同时,仅在一组预设的锚点上评估模型。在跳过的阶段,GeoShrink通过将最新观测到的创新的几何保留部分加到最近的精确输出来预测面向求解器的输出。我们从弦切向传输和往返线投影推导出这一规则,并建立了一个几何锚点间距原则,在固定覆盖范围和首跨度下最小化最大相邻间隙的扩展。该分析刻画了预测误差的几何闭合和传播,而不假设可以访问未来的模型输出。实验覆盖了图像、视频、运动和音频生成,以及适配的3D后端。在约5倍加速下,GeoShrink在FLUX PSNR上比最强的列出的基线提高了3.10 dB。在HunyuanVideo上,它实现了报告的4.99倍加速,并在ChronoMagic-Bench-150 PSNR上比最强的列出的保真度基线提高了5.44 dB。在固定评估预算下的比较进一步显示了在运动、音频、音乐和3D生成上的显著收益。

英文摘要

Diffusion transformers incur substantial inference cost through repeated model evaluations along a sampling trajectory. We introduce GeoShrink, a training-free acceleration method that retains the original solver grid while evaluating the model only at a prescribed set of anchors. At skipped stages, GeoShrink predicts the solver-facing output by adding a geometrically retained fraction of the latest observed innovation to the most recent exact output. We derive this rule from chordal tangent transport and round-trip line projection, and establish a geometric anchor-spacing principle that minimizes the largest adjacent gap expansion under fixed coverage and first span. The analysis characterizes the geometric closure and propagation of prediction errors without assuming access to future model outputs. Experiments cover image, video, motion, and audio generation, together with adapted 3D backends. At approximately $5\times$ acceleration, GeoShrink improves FLUX PSNR by 3.10 dB over the strongest listed baseline. On HunyuanVideo, it achieves a reported $4.99\times$ speedup and improves ChronoMagic-Bench-150 PSNR by 5.44 dB over the strongest listed fidelity baseline. Comparisons at fixed evaluation budgets further show substantial gains on motion, audio, music, and 3D generation.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Griffith University(格里菲斯大学)
  • Data61, CSIRO(澳大利亚联邦科学与工业研究组织数据61中心)
  • Artificial Intelligence Lab(人工智能实验室)
  • Institute of Deep Perception Technology(深度感知技术研究所)
  • JITRI(江苏省产业技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑