arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2305.04461cs.CVcs.GR

局部注意力SDF扩散用于可控三维形状生成

Locally Attentional SDF Diffusion for Controllable 3D Shape Generation

Xin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong, Yang Liu, Heung-Yeung Shum

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出局部注意力SDF扩散框架,通过两阶段扩散模型和视图感知局部注意力机制,实现基于草图的可控三维形状生成,提升局部可控性和泛化性。

中文摘要 AI 辅助

尽管近期三维生成神经网络的快速发展极大地改善了三维形状生成,但普通用户创建三维形状并控制生成形状的局部几何仍然不便。为解决这些挑战,我们提出了一种基于扩散的三维生成框架——局部注意力SDF扩散,通过二维草图图像输入来建模合理的三维形状。我们的方法构建于两阶段扩散模型之上。第一阶段称为占用扩散,旨在生成低分辨率的占用场以近似形状外壳。第二阶段称为SDF扩散,在第一阶段确定的占用体素内合成高分辨率的有符号距离场,以提取精细几何。我们的模型由一种新颖的视图感知局部注意力机制赋能,用于图像条件下的形状生成,该机制利用二维图像块特征来引导三维体素特征学习,极大地提高了局部可控性和模型泛化能力。通过在草图条件和类别条件下的三维形状生成任务中进行大量实验,我们验证并展示了我们的方法提供合理且多样化的三维形状的能力,以及其相对于现有工作的优越可控性和泛化性。我们的代码和训练模型可在https://zhengxinyang.github.io/projects/LAS-Diffusion.html获取。

英文摘要

Although the recent rapid evolution of 3D generative neural networks greatly improves 3D shape generation, it is still not convenient for ordinary users to create 3D shapes and control the local geometry of generated shapes. To address these challenges, we propose a diffusion-based 3D generation framework -- locally attentional SDF diffusion, to model plausible 3D shapes, via 2D sketch image input. Our method is built on a two-stage diffusion model. The first stage, named occupancy-diffusion, aims to generate a low-resolution occupancy field to approximate the shape shell. The second stage, named SDF-diffusion, synthesizes a high-resolution signed distance field within the occupied voxels determined by the first stage to extract fine geometry. Our model is empowered by a novel view-aware local attention mechanism for image-conditioned shape generation, which takes advantage of 2D image patch features to guide 3D voxel feature learning, greatly improving local controllability and model generalizability. Through extensive experiments in sketch-conditioned and category-conditioned 3D shape generation tasks, we validate and demonstrate the ability of our method to provide plausible and diverse 3D shapes, as well as its superior controllability and generalizability over existing work. Our code and trained models are available at https://zhengxinyang.github.io/projects/LAS-Diffusion.html

发表机构

  • Tsinghua University(清华大学)
  • Microsoft Research Asia(微软亚洲研究院)
  • Peking University(北京大学)
  • International Digital Economy Academy(国际数字经济研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑