arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11685cs.CVcs.AI

HI3D 3.0(Twinkle3D):面向特定物体的高分辨率3D资产生成

HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution

发表机构hi3d.ai
查看机构详情
  • hi3d.ai

机构由 AI 辅助整理,请以论文原文为准。

Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiw… 展开作者

Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiwen Wang, Jiaqi Wu, Kang Wu, Kaiqiao Yang, Zherui Yang, Hu Zhang, Xuezhi Zhao, Xinhe Zheng, Yukun Li, Heliang Zheng, Rongfei Jia

首次发表
浏览论文内容

中文总结 AI 辅助

Hi3D 3.0(Twinkle3D)是一款高保真图像到3D生成系统,通过优化DiT架构等技术,在$2048^{3}$分辨率下生成水密网格,在铭文恢复等指标上优于四款商业系统。

中文摘要 AI 辅助

图像到3D生成技术已能生成与输入图像高度相似的物体,而重现被描绘物体本身(包括定义其身份的特定几何结构)仍是一项重大挑战。铭文、品牌标识和重复结构对物体身份至关重要,但在现有技术中常出现扭曲或丢失的情况。本文提出Hi3D 3.0,这是一款针对物体特定保真度的图像到3D生成系统,其几何模型为Twinkle3D,可生成分辨率达$2048^{3}$的水密三角形网格。Twinkle3D在四个维度推进高保真几何生成:其一,O-Voxel/FaithC虽具备高表示精度,但常存在表面质量差、几何结构非水密的问题,我们在保留其$2048^{3}$级精度的同时解决了这两个问题;其二,我们通过重新设计的DiT架构和大规模分布式训练优化,将扩散生成扩展至最多30万个几何标记的序列,将每步训练时间从约10分钟缩短至10秒;其三,后续的细化无法完全弥补初始生成阶段引入的误差,因此我们强化初始生成阶段的全局形状和局部细节,最终得到的单阶段模型在$512^{3}$细化条件下,性能超越了此前的两阶段流水线;其四,我们引入细粒度图像-3D跨模态交互机制,强化视觉证据与几何标记的对应关系,提升物体特定结构的恢复能力。我们采用基于轮廓和法向量场的对齐指标评估几何保真度,Hi3D 3.0在所有报告指标上均优于四款商业系统,其铭文字符召回率达82.1%、精度达98.2%,而最强竞争对手的召回率仅为21.7%。

英文摘要

Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.

补充信息

↑