arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RenderBench:基于重建数字孪生的渲染到真实视频迁移基准

RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins

Dicong Qiu, Zhiyuan Xu, Yaosheng Liu, Feng Han, Bo Ye

arXiv 2610.08684首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Southeast University(香港科技大学(广州); 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RenderBench提出一个包含12个重建真实场景的基准,通过配对真实视频和可编辑数字孪生,评估渲染到真实视频迁移的外观保真度与几何动态保持。

AI 中文摘要

现代视频模型可以从真实外观参考和指定场景结构、视角变化及运动的代理渲染中生成逼真视频。评估这种渲染到真实(render-to-real)能力需要一段描绘相同场景演化的真实目标视频,以及一个可编辑、几何配准的3D副本。此类数据传统上需要大量的手动建模、标定和动画制作工作。我们引入了RenderBench,一个包含12个重建真实场景的基准,涵盖大规模室内环境和第一人称视角,并包含静态与动态设置。我们的构建流程结合了视觉几何、神经重建和辅助3D创作。每个场景被分解为静态物体和动态角色,配准到采集相机,并仅在通过多视图几何和时间验证后被接受。每个评估单元包含外观参考图像、留出的真实目标视频、可编辑数字孪生、匹配的代理渲染以及渲染器原生的场景标注。我们针对配对真实目标视频评估迁移模型,保留PAI-Bench-C兼容的结构投影,并使用场景标注按物体、可见性、关节和运动定位失败。首次发布保留了14个已配准样本中的12个(85.7%),包含1,496对配对真实-代理帧。所有发布的场景均通过文件完整性和环境编辑审计,而代理诊断产生深度si-RMSE为0.2170,实例mIoU为0.3673。RenderBench提供了配对真实观测和可编辑场景状态,用于评估外观保真度以及几何和动态的保持。

英文摘要

Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video depicting the same scene evolution, paired with an editable, geometrically registered 3D replica. Such data has traditionally required substantial manual modeling, calibration, and animation effort. We introduce RenderBench, a benchmark of 12 reconstructed real-world scenes spanning large-scale indoor environments and egocentric viewpoints, with both static and dynamic settings. Our construction pipeline combines visual geometry, neural reconstruction, and assisted 3D authoring. Each scene is decomposed into static objects and dynamic actors, registered to the capture cameras, and accepted only after multi-view geometric and temporal validation. Each evaluation unit contains appearance reference images, a held-out real target video, an editable digital twin, a matched proxy render, and renderer-native scene annotations. We evaluate transfer models against paired real target videos, retain PAI-Bench-C-compatible structural projections, and use scene annotations to localize failures by object, visibility, articulation, and motion. The first release retains 12 of 14 registered samples (85.7%), comprising 1,496 paired real-proxy frames. All released scenes pass file-integrity and environment-edit audits, while proxy diagnostics yield a depth si-RMSE of 0.2170 and instance mIoU of 0.3673. RenderBench provides paired real observations and editable scene state for assessing both appearance fidelity and preservation of geometry and dynamics.

Comments10 pages, 4 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑