arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HIVE-3D:用于高质量3D场景生成的分层体素增强

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

Bin Zang, Wenting Zheng, Xiaoliang Luo, Zhiyuan Fang, Shi Li, Lvchun Wang, Wei Yu, Yi Zhao, Tian Xie, Yuchi Huo, Rengan Xie

arXiv 2607.13468首次发表:更新:

AI 中文总结

研究针对单图像3D场景生成中分辨率受限问题,提出HIVE-3D方法。基于分层体素增强框架,经图像分割、注意力检索、构建分层组件树及体素超分辨率模型,实现粗到细的分层超分辨率,显著优于先前方法,达当前最优性能。

AI 中文摘要

最近,一系列作品能从单张图像生成令人印象深刻的3D对象,但受限于表示分辨率,不适用于3D场景生成。本文介绍HIVE-3D,一种基于分层体素增强框架的高质量3D场景生成新方法。输入单一场景图像,先生成粗糙初始场景,引入图像分割和基于注意力的检索使2D图像组件与3D场景组件对齐,组织场景关系成分层组件树,最后提出体素超分辨率模型生成精细体素。通过粗到细的分层超分辨率,生成高分辨率高质量3D场景,实验表明该方法显著优于先前方法。

英文摘要

Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce HIVE-3D, a novel method for high-quality 3D scene generation based on hierarchical voxel enhancement framework. Specifically, given a single scene image as input, we first produce a coarse initial scene, then introduce image segmentation and attention-based retrieval to align 2D image components with 3D scene components. Subsequently, we organize these scene relations into a hierarchical component tree, where nodes closer to the leaves denote finer-grained components. Finally, we propose a voxel super-resolution model that generates refined voxels for the target instance while maintaining strong consistency with the coarse voxels. Equipped with this model, we perform coarse-to-fine hierarchical super-resolution on images and voxels for each component, producing a high-resolution and high-quality 3D scene. Extensive experiments demonstrate that our method significantly outperforms previous approaches, achieving state-of-the-art performance.

CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026). Project page: https://xbdff.github.io/HIVE-3D/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑