发表机构
KRAFTON AI, Republic of Korea; Amazon, Australia(KRAFTON AI, 韩国; 亚马逊, 澳大利亚)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出首个针对图像到形状扩散Transformer的压缩方法,通过活力引导的结构化剪枝、自适应量化和微调,在保持几何保真度下实现高达66%的模型尺寸缩减。
AI 中文摘要
我们提出了首个针对图像到形状扩散Transformer(DiTs)的压缩方法,该方法在保持几何保真度的同时大幅减小模型尺寸。尽管3D形状生成取得了显著进展,但基于DiT的大型模型在资源受限环境中仍计算成本过高。此外,现有为不同领域开发的扩散模型压缩策略难以直接迁移到3D生成,而先前的3D效率方法主要关注推理速度而非主干网络压缩。为解决这一局限,我们构建了一个针对图像到形状DiT的几何感知压缩框架。基于3D DiT层对几何合成具有非均匀重要性的观察,我们引入了一个活力引导框架,整合了结构化剪枝、自适应量化和针对性微调。我们的方法在多种最先进的图像到3D模型上实现了高达66%的模型尺寸缩减,同时保持与全尺寸模型相当的合成保真度。这突显了我们的框架作为即插即用解决方案在多种模型上实现高效3D形状生成的潜力。
英文摘要
We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geometric fidelity. Despite remarkable progress in 3D shape generation, large DiT-based models remain computationally prohibitive in resource-constrained settings. Furthermore, it is difficult to directly transfer existing diffusion model compression strategies developed for different domains to 3D generation, and prior 3D efficiency approaches focus primarily on inference speed rather than backbone compression. To address this limitation, we build a geometry-aware compression framework tailored to image-to-shape DiTs. Guided by the observation that 3D DiT layers exhibit non-uniform importance for geometry synthesis, we introduce a vitality-guided framework integrating structured pruning, adaptive quantization, and targeted fine-tuning. Our method achieves up to 66% model-size reduction across state-of-the-art image-to-3D models while maintaining synthesis fidelity comparable to full-sized counterparts. This highlights the potential of our framework as a plug-and-play solution for efficient 3D shape generation across diverse models.
CommentsAccepted to ECCV 2026