Cyc3D:评估图像到3D生成中的循环结构稳定性与资产可用性
Cyc3D: Evaluating Cyclic Structural Stability and Asset Usability in Image-to-3D Generation
AI总结:
该研究针对图像到3D生成的评估缺陷,提出多维基准Cyc3D,通过多维度诊断揭示模型缺陷,实验发现闭源前馈模型在多项指标上优于开源优化基线,但最强方法仍存循环稳定性差距。
AI中文摘要:
图像条件3D生成已取得快速进展,但现有评估协议大多仅判断渲染视图的合理性与语义对齐性,却忽略了生成器是否形成稳定的3D解释,以及生成的资产是否可在图形管线中使用。我们提出Cyc3D,这是一个多维基准,沿两个互补轴评估图像到3D生成:跨视图对象一致性与表示质量。在资产层面,Cyc3D衡量对象身份在不同渲染视角间是否保持语义一致性;在模型层面,我们提出视图循环结构一致性,这是一种闭环渲染-再生-对齐协议,可从新视角反复重新观测生成的资产,并量化各代之间的几何、感知与语义漂移。为评估超越渲染外观的原生资产可用性,Cyc3D还评估几何结构、参考图像保真度、网格离散化与效率,以及UV参数化质量。这些诊断共同揭示了单一感知分数所掩盖的缺陷,并提供了模型不稳定性与表示缺陷的可解释证据。对5种代表性图像到3D系统的实验表明,闭源前馈模型在几何保真度、网格质量与循环稳定性上始终优于开源优化型基线。不过,即便是最强的方法,其循环稳定性得分也低于48,这表明在视觉上合理的生成与稳健的3D对象理解之间仍存在持续差距。
英文摘要:
Image-conditioned 3D generation has advanced rapidly, yet existing evaluation protocols largely judge rendered-view plausibility and semantic alignment, overlooking whether a generator forms a stable 3D interpretation and produces assets usable in graphics pipelines. We introduce Cyc3D, a multidimensional benchmark that evaluates image-to-3D generation along two complementary axes: Cross-View Object Consistency and Representation Quality. At the asset level, Cyc3D measures whether object identity remains semantically coherent across rendered viewpoints. At the model level, we propose View-Cycle Structural Consistency, a closed-loop render-regenerate-align protocol that repeatedly re-observes a generated asset from novel views and quantifies geometric, perceptual, and semantic drift across generations. To assess native asset usability beyond rendered appearance, Cyc3D further evaluates geometric structure, reference-image fidelity, mesh discretization and efficiency, and UV parameterization quality. Together, these diagnostics expose failures obscured by a single perceptual score and provide interpretable evidence of both model instability and representation defects. Experiments on five representative image-to-3D systems show that closed-source feed-forward models consistently outperform open-source optimization-based baselines in geometric fidelity, mesh quality, and cycle stability. Nevertheless, even the strongest methods achieve cycle-stability scores below 48, revealing a persistent gap between visually plausible generation and robust 3D object understanding.