arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28731cs.CGcs.GR

CuACD:一种完全GPU驻留的近似凸分解方法

CuACD: A Fully GPU-Resident Approximate Convex Decomposition

  • UC San Diego(加州大学圣地亚哥分校)
  • Sudo GmbH(Sudo有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Ruoxi Shi, Xinyue Wei, Fanbo Xiang, Zexiang Xu, Hao Su

AI总结:

针对现有ACD方法在CPU上运行缓慢的问题,提出CuACD,首个完全GPU驻留的近似凸分解系统,通过warp级融合与设备端堆分配器实现超一个数量级加速。

AI中文摘要:

近似凸分解(ACD)将三角形网格转换为小规模的凸部件集合,是物理仿真、碰撞检测和大规模机器人学习的标准预处理步骤。现代大多数ACD方法通过对候选切割平面进行昂贵的搜索来生成高质量分解,每个网格的运行时间长达数十秒,这迫使游戏管线进行过夜烘焙,并使得关节物体数据集在CPU集群上处理数天。先前的工作加速了孤立阶段,最近如VisACD基于GPU的可见性度量,但主要开销——搜索、网格切割和凸包构建——仍然留在CPU上,因为这些阶段自然分解为许多小规模同质阶段,在每个内核边界处逐渐减弱为最后的波次,并且每个阶段的可变大小输出迫使主机往返仅为了分配下一个启动的输入。我们通过采用warp(而非线程或线程块)作为算法设计单元来应对这些障碍,这一想法在图形处理社区中为不同的病理引入,我们在此将其改编为将计算几何管线的许多异质阶段融合为单个warp驻留内核,并配以设备端堆分配器,使得融合阶段之间的缓冲区可以在设备上调整大小和分配。基于此模板,我们提出了CuACD(CUDA ACD),第一个完全GPU驻留的ACD系统,以及一套可复用的GPU组件,作为开源独立CUDA模块发布,可插入任何基于搜索的ACD管线。在V-HACD基准、PartNet-Mobility和Objaverse子集上,CuACD在匹配或更优质量下比CoACD实现了超过一个数量级的加速。

英文摘要:

Approximate convex decomposition (ACD) converts triangle meshes into small sets of convex parts and is a standard preprocessing step for physics simulation, collision detection, and large-scale robot learning. The majority of modern ACD methods produce high-quality decompositions through an expensive search over candidate cutting planes, with per-mesh runtimes of tens of seconds that force game pipelines into overnight bakes and keep articulated-object datasets on CPU clusters for days. Prior work has accelerated isolated stages, most recently VisACD's GPU-based visibility metric, yet the dominant costs -- search, mesh cutting, and convex hull construction -- have remained on the CPU because their natural decomposition into many small homogeneous phases trails off in a fading last wave at every kernel boundary, and the variable-sized output of each phase forces a host round trip simply to allocate the next launch's input. We address these obstacles by adopting the warp, rather than the thread or thread block, as the unit of algorithm design, an idea introduced in the graph-processing community for a different pathology and which we adapt here to fuse the many heterogeneous phases of a computational-geometry pipeline into single warp-resident kernels, paired with a device-side heap allocator that lets the buffers between fused phases be sized and allocated on the device. Building on this template, we present CuACD (CUDA ACD), the first fully GPU-resident ACD system, together with a suite of reusable GPU components, released as open-source standalone CUDA modules that drop into any search-based ACD pipeline. On the V-HACD benchmark, PartNet-Mobility, and an Objaverse subset, CuACD achieves more than an order of magnitude of speedup over CoACD at matched or better quality.

补充信息

↑