arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23974cs.AI

LEAP-NBV:面向基础模型下一最佳视角规划的轻量级边缘主动感知

LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning

  • Johns Hopkins University(约翰霍普金斯大学)
  • US Army Research Laboratory(美国陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

Boxun Hu, Jiawei Ge, Axel Krieger, Peng Wang, Tinoosh Mohsenin

AI总结:

LEAP-NBV提出轻量级主动感知框架,通过蒸馏和量化在边缘设备上高效运行基础模型驱动的下一最佳视角规划,实现2倍加速和3倍能耗降低,同时保持任务质量。

AI中文摘要:

基础模型正赋予自主系统更高的智能,使其能够通过视觉感知更全面地理解环境。一个代表性例子是人体网格恢复(HMR),它提供目标3D姿态和形状的有用估计,可惠及战术任务。然而,此类模型的规模和功耗需求使其难以在边缘平台上运行,并限制了其实时性能,从而削弱了战术边缘部署的要求——尤其是在主动感知场景中,移动机器人必须在机载上规划下一最佳视角,且无法在通信受限的情况下卸载计算。我们提出LEAP-NBV,一种轻量级主动感知框架,可在边缘设备上机载运行基础模型驱动的下一最佳视角(NBV)规划。为此,我们将一系列大型HMR教师模型蒸馏为紧凑的32M学生模型,采用离线网格目标,然后将视觉编码器量化至FP16,并表征其在设备上的精度和延迟。在遮挡感知的主动感知循环中,我们在相同的留出基准上评估所有配置,并在NVIDIA Jetson Xavier NX上部署端到端流水线,报告实测的机载延迟和能耗。蒸馏在测试集上将未蒸馏学生的Procrustes对齐平均每顶点位置误差(PA-MPVPE)恢复6-7毫米。选择边缘最优压缩模型使HMR引擎达到约12毫秒,且精度损失较小,并以3.6 FPS和每帧2.6焦耳运行完整闭环,相比未压缩模型实现2.0倍加速和3.0倍能耗降低,同时下游任务质量几乎匹配。

英文摘要:

Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human Mesh Recovery (HMR), which provides useful estimates of a target's 3D pose and shape that can benefit tactical missions. However, the size and power demands of such models make them difficult to run on edge platforms and limit their real-time performance, undermining the requirements of tactical edge deployment - especially for active perception, where a mobile robot must plan its next-best view on-board and cannot offload computation under contested communications. We present LEAP-NBV, a lightweight active-perception framework that runs foundation-model-driven Next-Best-View (NBV) planning on-board an edge device. To this end, we distill a family of large HMR teachers, each into a compact 32M student, with an offline mesh objective, then quantize the vision encoder to FP16 and characterize its on-device accuracy and latency. Within an occlusion-aware active perception loop, we evaluate all configurations on the same held-out benchmark and deploy the end-to-end pipeline on an NVIDIA Jetson Xavier NX, reporting measured on-device latency and energy. Distillation recovers 6-7 mm of Procrustes-aligned mean per-vertex position error (PA-MPVPE) over the undistilled student on the test set. Selecting the edge-optimal compression model brings the HMR engine to ~12 ms at a small accuracy cost and runs the full closed loop at 3.6 FPS and 2.6 J per frame, achieving a 2.0x speedup and 3.0x lower energy than the uncompressed model while nearly matching downstream task quality.

↑