GeoCache:通过几何增量传输实现多视图纹理扩散的无训练加速
GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport
浏览论文内容
中文总结 AI 辅助
本研究提出无训练插件GeoCache,通过传输锚点视图的几何对齐更新实现多视图纹理扩散加速,在三个基准数据集上实现优于时间缓存和步骤缩减的速度-保真度权衡,2倍以上加速时表现突出。
中文摘要 AI 辅助
几何条件多视图扩散可实现高质量的3D纹理生成,但其每视图去噪器的重复评估会带来大量计算成本。现有的无训练加速器主要通过在去噪步骤间复用计算来利用时间冗余。然而在多视图纹理处理中,跳过一步也会消除跨视图交互,而这种交互能持续对齐同一表面的不同观测,进而导致一致性和保真度快速下降。我们的分析发现了一种互补的冗余来源:尽管中间特征仍为视图特定的,但几何对应的表面点在其预测的干净信号中表现出可迁移的演化。基于该观察,我们引入了GeoCache(一种无训练插件),它会评估旋转后的锚点视图子集,并将其与几何对齐的每一步xz更新传输到剩余视图。周期性的全视图计算可控制累积误差,而与采样器一致的重建则能保留去噪轨迹。GeoCache无需重新训练或修改架构,且使用几何条件纹理处理流程中已有的位置图。在Hunyuan3D-2.1、SyncMVD和MVPainter上,GeoCache在2倍以上操作点的速度-保真度权衡优于时间缓存和步骤缩减。在Hunyuan3D-2.1上,它实现了2.21倍的去噪器循环加速,MV-LPIPS为0.0293,MV-PSNR为33.60 dB,是所有2倍以上测试方法中保真度最佳的。相同的传输配置在SyncMVD上达到最高加速和最低FLOPs,而在MVPainter上,GeoCache在加速方法中实现了最低FLOPs和最佳保真度。这些结果表明,跨视图几何是多视图纹理扩散的有效加速轴。
英文摘要
Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy by reusing computation across denoising steps. In multi-view texturing, however, skipping a step also removes the cross-view interaction that continually aligns different observations of the same surface, leading to rapidly degraded consistency and fidelity. Our analysis identifies a complementary source of redundancy: although intermediate features remain view-specific, geometrically corresponding surface points exhibit transferable evolution in their predicted clean signals. Based on this observation, we introduce \gc{}, a training-free plugin that evaluates a rotating subset of anchor views and transports their geometry-aligned per-step $\xz$ updates to the remaining views. Periodic full-view computation controls accumulated error, while sampler-consistent reconstruction preserves the denoising trajectory. \gc{} requires neither retraining nor architectural modification and uses the position maps already available in geometry-conditioned texturing pipelines. Across Hunyuan3D-2.1, SyncMVD, and MVPainter, \gc{} achieves a stronger speed--fidelity trade-off than temporal caches and step reduction at operating points above $2\times$. On Hunyuan3D-2.1, it delivers a $2.21\times$ denoiser-loop speedup with an MV-LPIPS of 0.0293 and an MV-PSNR of 33.60 dB, providing the best fidelity among all tested methods above $2\times$. The same transferred configuration reaches the highest speedup and lowest FLOPs on SyncMVD, while \gc{} achieves the lowest FLOPs and best fidelity among the accelerated methods on MVPainter. These results establish cross-view geometry as an effective acceleration axis for multi-view texture diffusion.