arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关注可行视角:遮挡场景下的CAD重建主动视角选择

Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion

Kartik Bali, Mahish Guru, Yiderigun Borjigin, Alexandra Starostina, Christian J. Cyron, Roland Aydin

arXiv 2610.11954首次发表:更新:

发表机构

Helmholtz-Zentrum Hereon; Leuphana Universität Lüneburg; Universität des Saarlandes; Technische Universität Hamburg(亥姆霍兹中心赫伦研究所; 吕讷堡大学; 萨尔大学; 汉堡工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SightCAD框架,将视角可行性作为约束,联合训练视角选择器与CAD生成VLM,在多基准的遮挡场景CAD重建任务中显著优于基线,可执行程序占比高。

AI 中文摘要

CAD重建方法假设了一个现实中很少能实现的理想条件:可以从任意期望角度拍摄,获得不受限制的物体视觉信息。然而真实物体嵌入在场景中,比如固定在墙壁上、卡在角落、放置在地面上,这使得大部分视角球不可达,剩余视角的信息价值也不均等。我们提出了SightCAD,一种将视角可行性作为首要约束的参数化CAD重建框架。本研究中,我们考虑来自标准CAD基准的物体,将其嵌入具有离散视角球上物理推导可见性约束的真实室内场景中。一个学习得到的视角选择器必须为视觉-语言模型(VLM)选择K个可行视角,该VLM会生成可执行的CadQuery代码,代码的评分依据是执行后实体的几何保真度。由于奖励仅在离散视角选择、自回归生成和CAD内核执行后才会产生,我们提出了一种联合训练范式,让视角选择器和CAD生成VLM针对该奖励共同训练。学习到的选择策略与随机、均匀及覆盖贪婪方法有显著差异,在K∈{1,…,5}的预算下,其性能比表面积最大化(SA-max)方法高出最多6.4个交并比(IoU)点。该完整系统在DeepCAD和Fusion360物体的场景嵌入遮挡多视图渲染上,分别比最佳基线高出21和17个有效平均交并比(effective-mIoU)点;在工业T-LESS基准的测试时域规范化真实图像,以及MP6D工业金属零件基准的合成和真实图像上均表现优异,同时生成可执行程序的比例为所有方法中最高,无效代码率≤1.5%。

英文摘要

CAD reconstruction methods assume a luxury reality rarely grants: unrestricted visual access to the object, photographed from any desired angle. Real objects, however, are scene-embedded, bolted against walls, wedged into corners, resting on floors, where the scene renders much of the view sphere unreachable and the remaining views unequally informative. We introduce \textbf{SightCAD}, a framework for parametric CAD reconstruction that treats view feasibility as a first-class constraint. In this work we consider objects from standard CAD benchmarks embedded in realistic indoor scenes with physically derived visibility constraints over a discrete view sphere. A learned view selector must choose $K$ feasible views for a vision--language model (VLM) that generates executable CadQuery code, scored by geometric fidelity of the executed solid. Because reward arrives only after discrete view selection, autoregressive generation, and CAD-kernel execution, we propose a joint training paradigm in which the view selector and the CAD-generation VLM are trained together against this reward. The learned selection policy departs sharply from random, uniform, and coverage-greedy alternatives, outperforming surface-area maximization (SA-max) by up to $6.4$ Intersection-over-Union (IoU) points across budgets $K\in\{1,\dots,5\}$. The full system surpasses strong external baselines on scene-embedded, occluded multi-view renders of DeepCAD and Fusion360 objects ($+21$ and $+17$ effective-mIoU points over the best baseline, respectively), as well as on test-time domain-canonicalized real images from the industrial T-LESS benchmark and on both synthetic and real images from the MP6D industrial metal-parts benchmark, while producing the highest rate of executable programs of any method compared (invalid-code rate ${\leq}1.5\%$).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑