arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08891cs.CE

Ortho2CAD:使用视觉语言模型从正投影图生成3D CAD

Ortho2CAD: 3D CAD generation from orthographic drawings using vision language models

Aditya Joglekar, Amit Regmi, Kenji Shimada, Levent Burak Kara

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在解决工程设计中从正投影图到3D CAD模型转换的问题,核心方法是利用视觉语言模型Ortho2CAD,通过监督微调与强化学习训练,借助绘图生成器实现大规模学习,主要贡献是代码有效率高且3D CAD精度超基线。

中文摘要 AI 辅助

工程设计意图常通过光栅化正投影图传达,但下游工作流程需要可编辑和参数化定义的3D计算机辅助设计(CAD)模型。为此引入Ortho2CAD,一种视觉语言模型,能将光栅化正投影图直接转换为可编辑的CadQuery代码并进而转换为3D CAD模型。训练时,有明确标签用监督微调,无真实标签则用几何基础强化学习。创建基于pythonOCC的绘图生成器实现大规模学习。在现有数据集上,模型生成的代码语法有效率达100%,3D CAD交并比精度超所有基线,平均相对提升超7%。表明利用视觉语言模型及相关技术能有效推动正投影图到3D CAD的重建。

英文摘要

Engineering design intent is often communicated through rasterized orthographic drawings. However, downstream workflows require editable and parametrically defined 3D computer-aided design (CAD) models. To bridge this gap, we introduce vision language model (VLM) frameworks specifically designed to translate rasterized orthographic drawings into editable CadQuery code, which can then be converted into 3D CAD models. Firstly, due to unavailability of large scale orthographic drawing datasets, we create a pythonOCC-based drawing generator that renders first-angle orthographic projections from STEP models, with dashed hidden lines and bounding box dimensions, and generate over 1 million drawings from existing 3D CAD model datasets. We also create a dataset of 100 drawings with manually dimensioned features. We show that supervised fine-tuning applied on small open-source VLMs when paired CadQuery code is available improves reconstruction accuracy on corresponding test sets. For datasets without code labels, geometry-grounded reinforcement learning is performed which uses generated-solid intersection-over-union (IoU) with ground truth solid as the reward, improving code validity and cross-dataset generalization. Then, an inference time self-refinement framework for frontier VLMs is introduced which repeatedly repairs invalid codes and compares orthographic projections of generated 3D models with the input drawing to revise the CadQuery code. Our self-refinement framework with GPT 5.5 achieves 100% valid code generation and the highest mean IoU across all test sets, with an average relative improvement of more than 11% over the next-best method. We show that leveraging VLMs can effectively pave the way forward for orthographic drawing to 3D CAD reconstruction. Our implementation is available at https://github.com/AdityaJoglekar/Ortho2CAD.

↑