arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

流匹配中的主时间步稀疏矩阵分解受限初始化

Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching

Jiayang Gu, Zheng Fang, Lichaun Xiang, Fanghui Liu, Xu Cai, Hongkai Wen

arXiv 2609.15643首次发表:更新:

发表机构

University of Warwick; Bytedance; Apple Inc.(华威大学; 字节跳动; 苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对流匹配扩散模型微调中LoRA谱初始化失效的问题,提出Prism-LoRA,通过主时间步选择与主通道过滤改善梯度对齐,在多个基准上加速收敛并提升性能。

AI 中文摘要

流匹配扩散模型近来已成为高保真视觉生成的有力范式。然而,其高昂的微调成本限制了其在下游任务上的可扩展性。尽管低秩适应(LoRA)结合谱初始化在自回归语言模型中通过更好地对齐梯度方向,展现了加速收敛和性能提升,但我们发现该方法在扩散模型微调中未能带来类似收益,往往相比原始方法仅产生边际甚至负面的改进。我们将这一差异归因于LoRA的低秩参数化与流匹配目标所固有的高秩梯度之间的根本性不匹配。特别是,随机时间步采样在训练步骤间引入了方向异质的梯度信号,导致低秩更新下的梯度对齐不良。为解决此问题,我们提出了Prism-LoRA,一种通过稀疏矩阵分解进行主时间步受限初始化的框架,以改善微调期间的梯度对齐。我们的方法包含两个关键组件:(i)主时间步选择,将初始化梯度限制在主导时间步的子集上,以抑制有效梯度秩;(ii)主通道过滤,移除任务无关的通道,使一步谱初始化梯度能更好地与长期优化轨迹对齐。大量实验表明,我们的方法在多个扩散微调基准上(包括主体驱动生成、可控生成和去模糊)一致地提升了收敛速度和最终性能,不仅在性能上优于基线LoRA及其他谱初始化方法,而且实现了更早阶段的收敛。

英文摘要

Flow-matching diffusion models have recently emerged as a strong paradigm for high-fidelity visual generation. However, their prohibitively high fine-tuning cost limits scalability to downstream tasks. While Low-Rank Adaptation (LoRA) combined with spectral initialization has demonstrated accelerated convergence and improved performance in autoregressive language models by better aligning gradient directions, we find that it fails to deliver similar gains in diffusion fine-tuning, often yielding marginal or even negative improvements over vanilla LoRA.We attribute this discrepancy to a fundamental mismatch between LoRA's low-rank parameterization and the intrinsically high-rank gradients induced by the flow-matching objective. In particular, stochastic timestep sampling introduces directionally heterogeneous gradient signals across training steps, leading to misaligned updates under low-rank constraints.To address this issue, we propose Prism-LoRA,a Principal-timestep Restricted Init via Sparse Matrix-decomposition framework that improves gradient alignment during fine-tuning. Our method consists of two key components: (i) principal timestep selection, which restricts initialization gradients to a subset of dominant timesteps to suppress effective gradient rank, and (ii) principal channel filtering, which removes task-irrelevant channels, enabling the one-step spectral initialization gradient to better align with the long-horizon optimization trajectory. Extensive experiments demonstrate that our method consistently improves both convergence speed and final performance across multiple diffusion fine-tuning benchmarks, including subject-driven generation, controllable generation, and deblurring, achieving not only performance improvement but also earlier stages of convergence over baseline LoRA and other spectral-init methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑