arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RealCAD:面向存在域偏移与参数偏差的真实场景图像到CAD模型重建

RealCAD: Towards Real-World Image-to-CAD Reconstruction under Domain Shift and Parameter Bias

Yihe Sun, Ziyu Lu, Kaihua Tang, Xian-Sheng Hua

arXiv 2608.30617首次发表:更新:

发表机构

Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RealCAD是解决图像到CAD重建中域偏移与参数偏差的统一框架,通过多层面优化,在提升真实域性能的同时保持合成域竞争力,还构建了OpenRealCAD数据集。

AI 中文摘要

从图像重建可编辑的计算机辅助设计(CAD)模型对下游修改、制造和设计复用至关重要。然而,现有的图像到CAD方法主要基于合成渲染图开发,面临两个耦合的障碍:合成图像与真实图像之间存在显著的外观域差距,以及广泛使用的CAD数据中此前被忽视的参数偏差。我们表明,DeepCAD采用的局部归一化将多个几何参数集中在少数离散值附近,同时在单一尺度因子中编码了大量信息。因此,模型可以通过利用这些频繁出现的值而非从输入图像推断几何结构,实现看似很高的参数准确率。在本文中,我们提出了RealCAD,这是一个在表示、图像和特征层面解决这些限制的统一框架。在表示层面,我们将尺度信息重新分配到相应的几何参数,在共享尺度空间中产生更分散的参数分布。在图像层面,几何约束转换将合成渲染图转换为真实图像域,同时以对象轮廓为条件。在特征层面,多正对比目标对齐同一CAD模型在不同视角和图像域的表示,从而能够从每个单独的视角预测CAD命令序列。我们进一步引入了OpenRealCAD,包含392个3D打印对象的四视图照片及对应的真实命令序列。实验表明,修改后的表示大幅降低了从参数频率先验可获得的准确率,使参数准确率成为图像条件几何推断更可靠的度量。RealCAD进一步提高了真实域的命令和参数准确率,同时保持了具有竞争力的合成域性能。

英文摘要

Reconstructing editable Computer-Aided Design (CAD) models from images is essential for downstream modification, manufacturing, and design reuse. However, existing image-to-CAD methods are developed predominantly on synthetic renderings and face two coupled obstacles: a substantial appearance domain gap between synthetic and real images, and a previously overlooked parameter bias in widely used CAD data. We show that the local normalization adopted by DeepCAD concentrates several geometric parameters around a few discrete values while encoding substantial information in a single scale factor. Consequently, a model can achieve deceptively high parameter accuracy by exploiting these frequent values rather than inferring geometry from the input image. In this paper, we propose RealCAD, a unified framework that addresses these limitations at the representation, image, and feature levels. At the representation level, we redistribute scale information to the corresponding geometric parameters, producing less concentrated parameter distributions in a shared scale space. At the image level, geometry-constrained translation converts synthetic renderings toward the real-image domain while conditioning on object contours. At the feature level, a multi-positive contrastive objective aligns representations of the same CAD model across viewpoints and image domains, enabling CAD sequence prediction from each individual view. We further introduce OpenRealCAD, comprising four-view photographs of 392 3D-printed objects paired with ground-truth command sequences. Experiments show that the revised representation substantially reduces the accuracy attainable from parameter-frequency priors, making parameter accuracy a more reliable measure of image-conditioned geometric inference. RealCAD further improves real-domain command and parameter accuracy, while retaining competitive synthetic-domain performance.

CommentsThe code and dataset are publicly available. Code: https://github.com/sunyh39/RealCAD. Dataset: https://www.modelscope.cn/datasets/yeguomao/RealCAD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑