arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05539cs.CV

OmniMech:面向3D重建的全合一多模态机械基准

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Taiting Lu, Runze Liu, Ziwei Dong, Sisong Bei, Jingying Zeng, Mingjia Wang, Zhenghao Li, Kaiyuan Lin, Yi-Shan Wu, Yangshoudu Zheng, Hongxing Pan, Kai Zhang, Guo… 展开作者

Taiting Lu, Runze Liu, Ziwei Dong, Sisong Bei, Jingying Zeng, Mingjia Wang, Zhenghao Li, Kaiyuan Lin, Yi-Shan Wu, Yangshoudu Zheng, Hongxing Pan, Kai Zhang, Guoliang Shi, Ling Ma, Yifan Yang, Jiaying Lu, Qi He, Sung-Liang Chen, Yi-Chao Chen, Yincheng Jin, Mahanth Gowda

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员提出OmniMech——首个百万级多模态机械基准,针对工业制造数据的可执行CAD生成任务,含四类任务,实验发现现有模型在相关任务中存在不足,将发布资源支持后续研究。

中文摘要 AI 辅助

近期的视觉-语言模型(VLMs)可从图像生成可执行CAD程序,但现有方法主要针对粗糙的通用3D对象,极少涉及工业机械设计所需的精细几何与毫米级公差。我们提出OmniMech,首个用于评估VLMs从工业制造数据生成可执行CAD程序的百万级基准。OmniMech包含超过251000份带完整尺寸标注与公差的二维正投影视图,搭配原生CAD模型、多视图渲染、网格、STEP及B-rep表示,以及丰富语义标注。该基准涵盖四项任务:(1)从工程图生成参数化CAD程序;(2)图转3D以实现几何与结构一致的重建;(3)基于尺寸、符号、特征标注及制造约束的标注导向推理;(4)利用可视化、测量、CAD执行与验证工具的工具增强智能体推理。实验显示,当前VLMs与CAD专用模型在可执行程序合成、精细3D重建及可靠执行尺寸与公差方面仍存在困难。我们将发布基准数据、评估代码及工具接口以支持未来研究。

英文摘要

Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.

↑