arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12442cs.CV

LEGO:一种用于异中心视角到自我中心视角视频生成的无提升方法

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park, Taegyu Lim

首次发表
浏览论文内容

中文总结 AI 辅助

LEGO提出无提升的LVSM风格Transformer视角合成器,解决异中心到自我中心视频生成的跨视角对应问题,性能优于现有方法且可泛化至其他数据集。

中文摘要 AI 辅助

从单个异中心视角(第三人称视角)录制生成自我中心视角(第一人称视角)视频是新颖视角合成中极具挑战性的案例,因为两种视角的相机重叠度极低,目标视角的大部分区域未被观测到。当前最先进的方法通过估计深度、将视频提升为点云并从自我中心视角重新渲染来显式重建场景,以此条件视频扩散模型。这种确定性映射将每个像素分配到单个重投影位置,虽保留了纹理,但会将深度误差转化为内容错位。本文探究视频扩散模型应接收何种条件,提出一种无提升的解决方案:学习型视角合成器,即经微调的LVSM风格Transformer,无需深度、点云或重投影即可直接渲染自我中心视角,在内部解决跨视角对应关系。相比之下,其概率映射会根据学习到的对应分布,对每个区域在候选源位置上取平均,在保留结构的同时精细纹理会被平均消除。本文认为这种权衡适配扩散生成器,其去噪训练擅长恢复细节,因此有效条件应优先考虑结构对齐而非清晰度。该分布的集中度还会产生每个区域的置信度,用于在早期布局形成的去噪步骤中屏蔽低置信度区域,并引导生成器向高置信度区域发展。本文方法始终优于最先进的显式流水线,且无需重新训练即可泛化到其他数据集。合成器提供视角结构,扩散模型提供细节。

英文摘要

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.

发表机构

  • GenGenAI

机构由 AI 辅助整理,请以论文原文为准。

↑