Kaininja:将原生3D生成器扩展到部件级别
KaiNinja: Extending Native 3D Generators to the Part Level
- Alaya Lab(阿亚实验室)
- The University of Tokyo(东京大学)
- University of California, Merced(加州大学默塞德分校)
- Institute of Science Tokyo(东京科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对原生3D生成器只能输出融合整体、无法支持部件级操作的问题,提出基于双体积O-Voxel表示的Kaininja,扩展TRELLIS.2至部件级,无需分割网络,保持速度与质量,并显著提升整体保真度和部件精度。
中文摘要 AI 辅助
原生3D生成器将一张图像转换为单个网格。TRELLIS.2及其同类方法能够生成具有材质的高保真非水密几何体,但输出是一个融合的整体对象,而下游工作(如编辑、绑定和模拟)则基于部件级资产进行操作。一个朴素的想法是对TRELLIS.2生成的融合网格运行3D分割网络,但此类流程速度慢且受限于分割精度。我们希望找到一种简单的方法,将现有的原生3D生成器扩展到部件级别。但我们面临一个关键问题:O-Voxel网格在每个体素中仅存储一个表面片,因此单个体积无法表示两个部件接触的界面,无论分辨率如何。我们引入了一种双体积表示来解决此问题,并提出了KaiNinja,这是TRELLIS.2的部件级扩展,基于其O-Voxel表示的双体积形式构建。KaiNinja保持了TRELLIS.2的生成速度和质量,同时将其扩展到部件级别,且流程中无需掩码或分割器。其训练数据来自多种来源,包括CAD模型和由LLM驱动的智能体创作的资产;据我们所知,这是首个在智能体创作的部件数据上训练的3D生成模型。令人惊讶的是,我们还发现,与在同一数据集上微调的相同骨干网络相比,整体对象的保真度有所提高。与不同范式的部件生成流程相比,它将整体对象的倒角距离降低了40%,并将严格部件F分数提高了16%。
英文摘要
Native 3D generators turn one image into a single mesh. TRELLIS.2 and its peers deliver high-fidelity non-watertight geometry with materials, but the output is one fused object, while downstream work such as editing, rigging and simulation operates on part-level assets. A naive idea is to run a 3D segmentation network on the fused mesh that TRELLIS.2 generates, but such pipelines are slow and bounded by the accuracy of the segmentation. We want a simple way to extend an existing native 3D generator to the part level. But we face a critical problem: the O-Voxel grid stores one sheet of surface per voxel, so a single volume cannot represent the interface where two parts touch, at any resolution. We introduce a dual-volume representation to solve this problem and put forward KaiNinja, a part-level extension of TRELLIS.2 built on a dual-volume form of its O-Voxel representation. KaiNinja keeps the generation speed and quality of TRELLIS.2 while extending it to the part level, with no mask or segmenter in the pipeline. Its training data come from sources of many kinds, including CAD models and assets authored by an LLM-driven agent; to our knowledge it is the first 3D generative model trained on agent-authored part data. Surprisingly, we also find that whole-object fidelity improves over the same backbone fine-tuned on the same dataset. Against part generation pipelines of different paradigms, it lowers whole-object Chamfer distance by 40% and raises strict part F-score by 16%.