发表机构
Indian Institute of Technology Jodhpur(焦特布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
H3DNAS是一种无需源代码的ONNX原生3D点云模型压缩框架,通过三方面贡献在ModelNet40上实现多款3D点云模型参数减少与推理加速,精度损失极小。
AI 中文摘要
将3D点云模型部署到NVIDIA Jetson Orin Nano等边缘硬件时,会受到计算和内存预算的严重限制。现有的压缩方法需要访问模型的原始源代码,因此不适用于供应商和模型仓库通常分发的Open Neural Network Exchange(ONNX)二进制文件。本文提出H3DNAS,这是一种硬件感知的模型压缩框架,可直接在ONNX计算图上运行,无需在搜索过程中使用原始源代码、架构类定义或梯度访问。H3DNAS有三项贡献:(1)通道依赖图(Channel Dependency Graph, CDG),它将ONNX算子分为四类约束,并正式证明自由参数比例ρ_f是拓扑不变量,这是一个可在O(|V|+|E|)时间内计算的可证明压缩上限;(2)两阶段分层搜索,它通过L1重要性通道选择修剪候选架构,以输出保真度作为零样本无标签代理对其进行排名,并对帕累托最优候选应用GhostConv结构变异;(3)首个无需源代码的3D点云模型压缩流水线,完全通过ONNX图操作实现,无需原始架构定义。在ModelNet40上,H3DNAS将PointNet、PointNet++和PointMLP的参数数量分别减少65.5%、43.2%和49.1%,同时实现1.99倍、1.29倍和1.67倍的推理加速,且精度损失可忽略不计。源代码已公开。
英文摘要
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained by compute and memory budgets. Existing compression methods require access to the model's original source code, rendering them inapplicable to the Open Neural Network Exchange (ONNX) binaries commonly distributed by vendors and model repositories. We present \textbf{H3DNAS}, a hardware-aware model compression framework that operates directly on ONNX computational graphs without requiring original source code, architecture class definition, or gradient access during search. H3DNAS makes three contributions: (1) a \textbf{Channel Dependency Graph (CDG)} that classifies ONNX operators into four constraint classes and formally establishes that the free parameter fraction $ρ_f$ is topological invariant, a provable compression ceiling computable in $\mathcal{O}(|V|+|E|)$; (2) a \textbf{Two-Stage Hierarchical Search} that prunes candidate architectures by $L_1$-importance channel selection, ranks them by output fidelity as a zero-shot label-free proxy, and applies GhostConv structural mutation to Pareto-optimal candidates; and (3) the \textbf{first source-code-free compression pipeline for 3D point cloud models}, operating entirely via ONNX graph surgery with no original architecture definition required. On ModelNet40, H3DNAS reduces the number of parameters in PointNet, PointNet++, and PointMLP by $65.5\%$, $43.2\%$, and $49.1\%$, respectively, while achieving $1.99\times$, $1.29\times$, and $1.67\times$ inference speedups with negligible loss in accuracy. The source code is publicly available\footnote{https://github.com/ClarityLab-Org/h3dnas}.