arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

去噪扩散生成模型暗中计算注意力

Denoising Diffusion Generative Models Secretly Calculate Attentions

Farzan Haddadi, Leila Monfared, Ebrahim Rezaii, Mohammadreza Malek-Mohammadi, Pejman Zakalvand, Narges Mokhtari

arXiv 2609.00885首次发表:更新:

发表机构

School of Electrical Engineering, Iran University of Science & Technology(伊朗科技大学电气工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现去噪扩散模型内在使用与Transformer相似的注意力机制,基于此提出简化的注意力式扩散图像生成算法,可在少资源下达到相当性能,推动了注意力作为通用ML原理的认知。

AI 中文摘要

去噪扩散模型是图像生成领域的主导架构,而多数自然语言生成与建模主要由采用注意力机制的知名Transformer架构处理。本文表明,扩散模型也内在地使用与Transformer非常相似的注意力机制,因此注意力成为基于通用训练目标的通用机器学习原理。我们还展示了自编码器与基于注意力的模型在基本功能原理上的相似性,这些等价性使我们能根据实际需求在这些设计间进行互换。例如,我们可重构扩散框架以缩短冗长的训练过程并减少计算密集型的图像生成,采用该方法,我们提出了一种基于注意力机制的简化图像生成算法,结果表明,基于注意力的实现能达到相当的性能,且所需工作量和计算资源显著更少。

英文摘要

Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.

Commentssubmitted to IEEE Trans on Pattern Recog. Machine Intellig

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑