arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniCAD:面向机器人装配任务的3D空间推理大规模基准测试集

OmniCAD: A Large-Scale Benchmark for 3D Spatial Reasoning in Robotics Assemblies

Mingjia Wang, Taiting Lu, Ziwei Dong, Sisong Bei, Jingying Zeng, Runze Liu, Kaiyuan Lin, Hongxing Pan, Kai Zhang, Yizheng Hou, Yangshoudu Zheng, Chenchen Guo, Weiyuan Meng, Shubin Lyu, Zhijun Zheng, Dexu Wang, Xinyu Bai, Shurui Qian, Zhangzixin, Mengyu Pan, Guoliang Shi, Ling Ma, Yifan Yang, Qi He, Yi-Chao Chen, Yincheng Jin, Sung-Liang Chen, Mahanth Gowda

arXiv 2608.22637首次发表:更新:

AI 中文总结

本文提出OmniCAD——涵盖多类工业系统的3D空间推理大规模基准测试集,评估VLMs三类装配推理能力,发现当前VLMs在该任务中表现不佳,将开源相关资源推动该领域研究。

AI 中文摘要

近期的视觉语言模型(VLMs)在机器人感知与空间推理方面展现出较强能力,但其对复杂机械装配体的推理能力仍未得到充分探索。本文提出OmniCAD,这是一个面向各类工业系统的装配感知3D空间推理大规模基准测试集,涵盖机器人机构、汽车部件、航空航天结构及农业机械等领域。OmniCAD包含25000个机械装配体,每个装配体平均有12个零件,涉及21种配合关系。每个装配体均包含经人工验证的真实3D模型及20个视角的渲染图。该基准测试集评估三类能力:(1)零件级3D空间推理,要求预测零件的位置与姿态;(2)零件间关系推理,要求识别配合关系与装配约束;(3)工具增强型智能体推理,要求模型迭代选择视角、检查视觉证据并优化预测。实验表明,当前VLMs在工业装配推理中表现不佳,常生成不准确的姿态、无效的配合关系、零件穿透现象,且性能随装配复杂度提升而下降。我们将开源该基准测试集、评估代码及工具接口,以支持准确、物理有效且可扩展的3D装配推理研究。

英文摘要

Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex mechanical assemblies remains underexplored. We introduce OmniCAD, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery. OmniCAD contains 25k mechanical assemblies, with an average of 12 parts per assembly and 21 types of mate relationships. Each assembly includes a human-verified ground-truth 3D model and renderings from 20 viewpoints. The benchmark evaluates three capabilities: (1) component-level 3D spatial reasoning, requiring prediction of part positions and orientations; (2) part-to-part relational reasoning, requiring identification of mating relationships and assembly constraints; and (3) tool-augmented agentic reasoning, where models iteratively select viewpoints, inspect visual evidence, and refine predictions. Experiments show that current VLMs struggle with industrial assembly reasoning, often producing inaccurate poses, invalid mating relationships, part interpenetration, and degraded performance as assembly complexity increases. We will open-source the benchmark, evaluation code, and tool interfaces to support research on accurate, physically valid, and scalable 3D assembly reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑