arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于快速无训练球面全景图生成的线性融合多扩散

Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation

Akio Hayakawa, Yusuke Mukuta, Tatsuya Harada

arXiv 2609.01997首次发表:更新:

发表机构

The University of Tokyo; RIKEN Center for Advanced Intelligence Project(东京大学; 理化学研究所先进智能项目中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出LF-MultiDiffusion,一种扩展MultiDiffusion的无训练球面全景生成方法,通过重新表述潜在聚合问题提升效率与质量,实现15.36倍加速且性能优于基线。

AI 中文摘要

我们提出了LF-MultiDiffusion,这是一种无训练的全景图生成方法,它扩展了MultiDiffusion以支持目标与参考图像空间之间的线性投影。我们的核心思路是将潜在聚合重新表述为正则化最小二乘问题,并在去噪循环内使用基于Krylov的迭代求解器高效求解。该表述比以往的无训练方法能实现更密集、更自然的映射,用少得多的透视视图即可生成更稳定的结果。因此,LF-MultiDiffusion减少了去噪过程中图像生成器的评估次数,显著提升了推理效率。实验表明,LF-MultiDiffusion在视觉质量、文本对齐和全景一致性方面优于最强的无训练基线,同时实现了15.36倍的加速。我们的项目页面可访问此https URL。

英文摘要

We propose LF-MultiDiffusion, a training-free panorama generation method that extends MultiDiffusion to support linear projections between target and reference image spaces. Our key idea is to reformulate latent aggregation as a regularized least-squares problem and solve it efficiently with a Krylov-based iterative solver inside the denoising loop. This formulation enables denser and more natural mappings than prior training-free methods, yielding more stable generation with far fewer perspective views. As a result, LF-MultiDiffusion reduces the number of image generator evaluations during denoising and significantly improves inference efficiency. Experiments show that LF-MultiDiffusion achieves better visual quality, text alignment, and panoramic consistency than the strongest training-free baseline, while providing a 15.36$\times$ speedup. Our project page is available at: https://ahykw.github.io/lfmd.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑