arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32761cs.CV

从前馈到流:统一重建与生成比想象中更简单

From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think

Haoru Wang, Qianfan Shen, Kai Ye, Wenzheng Chen, Baoquan Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种统一的基于流的公式,使重建与生成共享同一预测器和主干网络,通过单步与多步去噪配置切换,在3D任务中同时提升重建质量与生成细节保真度。

中文摘要 AI 辅助

在图像提供证据的地方进行重建,在没有证据的地方进行生成:诸如Atlas(World Labs Team, 2026)等空间世界模型的最新成功凸显了在一个模型中统一重建与生成的价值。然而,这两者长期以来一直处于不同的范式之中,并具有各自独特的失败模式:前馈重建将模糊性平均化为模糊,而条件生成则发明出看似合理但场景不一致的细节。在这项工作中,我们提出了一种统一的基于流的重建与生成公式,其中共享的干净目标预测器在其单步端点执行直接重建,并通过多步流展开条件生成。一项受控的玩具研究揭示了其机制:在单步情况下,预测器会像前馈方法一样坍缩为条件均值,偏向一致性而非多样性。在多步推理中,生成细节的保真度随上下文丰富度而增长:更接近的观测减少了模糊性,并产生了更匹配的细节。我们进一步在表观和几何3D任务中实例化了该公式。JiT-LVSM在新视图合成中提高了感知和分布质量,而JUSt3R在保持具有竞争力的单步几何预测的同时,增加了多步推理能力,减少了面纱和飞点伪影,在更多测试时计算下产生更干净的表面结构。它们共同表明,重建和生成可以共享一个公式和一个主干网络,其行为由去噪配置控制——使得统一变得出奇地简单。

英文摘要

Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive failure modes: feed-forward reconstruction averages ambiguity into blur, while conditional generation invents plausible but scene-inconsistent detail. In this work, we present a unified flow-based formulation for reconstruction and generation, where a shared clean-target predictor performs direct reconstruction at its single-step endpoint and unfolds conditional generation through multi-step flow. A controlled toy study reveals the mechanism: with a single step, the predictor collapses to the conditional mean just like feed-forward methods, favoring consistency over diversity. With multi-step inference, the fidelity of generated details grows with context richness: closer observations reduce ambiguity and yield better-matched details. We further instantiate the formulation in appearance and geometry 3D tasks. JiT-LVSM improves perceptual and distributional quality in novel view synthesis, while JUSt3R retains competitive single-step geometry prediction with additional multi-step inference capabilities that reduces veil and flying-pixel artifacts, producing cleaner surface structure with greater test-time compute. Together, they show that reconstruction and generation can share both a formulation and a backbone, with their behavior governed by denoising configuration---making unification surprisingly simple.

补充信息

↑