arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35570cs.ROcs.CV

EdgeVLN:面向运行时感知的部署就绪量化视觉语言导航模型

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Rithvik Jonna, Man Namgung, Aakash Gurram, Tinoosh Mohsenin

首次发表
浏览论文内容

中文总结 AI 辅助

EdgeVLN通过量化StreamVLN与LATTE停止预测器,在边缘设备上实现部署就绪的VLN,4位模型SR达58.02%,内存、延迟和能耗显著优化。

中文摘要 AI 辅助

视觉语言导航(VLN)模型性能优异,但面向计算资源丰富的平台,限制了其在内存和功耗受限的机器人边缘设备上的部署。仅靠压缩并不能确定VLN模型在保持导航行为的同时,是否满足边缘平台的内存、延迟和能量预算。我们提出EdgeVLN,一种运行时感知、部署就绪的量化VLN模型,填补了这一空白。EdgeVLN结合了量化后的StreamVLN模型与潜在轨迹终止提取器(LATTE),后者是一个轻量级因果Transformer,通过预测停止动作验证器排名来改进实时停止。两者均通过我们的this http URL VLN驱动程序执行,该驱动在板载重建流式上下文并剪枝内存令牌。我们针对从8位到2位的权重量化以及多种推理运行时,对预训练的StreamVLN骨干进行了表征,以确定可行的运行点。LATTE在量化释放的预算内重用骨干隐藏状态,既不需要第二个视觉编码器,也不需要额外的骨干前向传播。我们在BF16和IQ4 NL格式下,对全部1,839个R2R VLN-CE验证未见环境中的六个骨干精度和七个候选停止头进行了评估。我们在NVIDIA Jetson Orin NX 16 GB上测量了模拟中的成功率(SR)以及延迟、能量和常驻内存。LATTE在部署的4位模型上实现了我们最高的SR,达到58.02%,超过了BF16基线,且每个导航步骤仅增加0.013秒延迟。四种位格式实现了几乎相同的SR,但每步能量因执行路径不同而相差36.8倍。在我们的VLN驱动程序下,仅IQ4 NL适合该板卡,使用11.35 GB常驻内存,同时比存储流式BF16快20.8倍,能耗低13.3倍。INT2格式崩溃。运行时选择、内存令牌剪枝和量化对于高效边缘部署至关重要。

英文摘要

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavior. We introduce EdgeVLN, a runtime-aware, deployment-ready quantized VLN model that closes this gap. EdgeVLN combines a quantized StreamVLN model with Latent Trajectory Termination Extractor (LATTE), a lightweight causal transformer that improves real-time stopping by predicting a Stop Action verifier rank. Both execute through our llama.cpp VLN driver, which reconstructs streaming context and prunes memory tokens on-board. We characterize a pretrained StreamVLN backbone across weight quantization from 8 to 2 bits and multiple inference runtimes to identify a feasible operating point. LATTE reuses backbone hidden states within the budget freed by quantization, requiring neither a second vision encoder nor an additional backbone forward pass. We evaluate six backbone precisions and seven candidate stop heads on BF16 and IQ4 NL across all 1,839 R2R VLN-CE val-unseen episodes. We measure success rate (SR) in simulation and latency, energy, and resident memory on an NVIDIA Jetson Orin NX 16 GB. LATTE achieves our highest SR, 58.02 percent on the deployed 4-bit model, exceeding the BF16 baseline with only 0.013 s additional latency per navigation step. Four-bit formats achieve nearly identical SR, but step energy varies 36.8 times by execution path. Only IQ4 NL under our VLN driver fits the board, using 11.35 GB resident memory while running 20.8 times faster and using 13.3 times less energy than storage-streamed BF16. INT2 collapses. Runtime selection, memory-token pruning, and quantization are essential for efficient edge deployment.

发表机构

  • Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑