arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

I2CD:面向仿真就绪碰撞几何的直接图像到凸分解

I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry

Qian Wang, Liam Merz Hoffmeister, Brian Scassellati, Daniel Rakita

arXiv 2610.03453首次发表:更新:

发表机构

Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

I2CD直接从单张RGB图像预测凸分解,冻结预训练扩散模型仅训练轻量交叉注意力头,生成无需后处理的凸几何,速度提升6-37倍,在多个仿真器和物理机器人上验证了高效性与可靠性。

AI 中文摘要

物理仿真器和运动规划器需要凸碰撞几何,然而图像到三维生成模型输出的是密集且通常非流形的视觉网格。目前,连接两者需要经过修复、抽取和近似凸分解的缓慢且脆弱的“重建-分解”流水线。我们提出I2CD,它直接从单张RGB图像预测凸分解。I2CD并非训练新的图像到三维模型,而是冻结预训练的Hunyuan3D-2图像条件扩散变换器和形状解码器,仅训练一个轻量级交叉注意力头(38M参数,不到十个GPU小时),其学习到的“凸槽”令牌发出K个凸多面体的半平面参数。输出紧凑、构造上保证凸性,无需任何后处理即可直接加载到物理引擎中,每张图像约0.5秒。在227个保留的OmniObject3D和Google Scanned Objects实例上,I2CD在八个“重建-分解”流水线中实现了最高的体积IoU,同时端到端速度快6到37倍。在MuJoCo、PyBullet、Genesis和Isaac Sim的跨仿真器研究中,每个引擎都按原样使用I2CD几何体,而原始生成的网格虽然“加载”成功,但在大多数情况下被静默替换为不同的碰撞形状,或者需要数秒到数分钟的逐对象预处理。在物理xArm7上,I2CD为包含20个物体的杂乱场景生成规划就绪几何体仅需11秒,而最强基线需要328秒,拾放执行成功率相当(100次试验中85次对90次)。

英文摘要

Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned "convex-slot" tokens emit the halfplane parameters of $K$ convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in ${\sim}0.5$s per image. On $227$ held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running $6$-$37\times$ faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes "load" everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a $20$-object cluttered scene in $11$s versus $328$s for the strongest baseline, at comparable pick-and-place execution success ($85$ vs. $90$ of $100$ trials).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑