发表机构
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Scout系统通过间歇调用云端VLM为边缘模型增量学习新类别,在NVIDIA Jetson Orin Nano上实现接近预定义列表的准确率,能耗降低59%-71%。
AI 中文摘要
大型视觉语言模型(VLMs)能够实现超越固定类别集合的识别,但其计算需求使得它们无法在许多边缘设备上运行。云端卸载使这一能力变得可用,但发送每张图像会消耗稀缺的带宽和通信能量。我们探讨如何在严格的计算、能量和带宽预算内,将VLM的开放世界识别能力带到边缘端。野生动物监测为探索这一问题提供了自然场景,因为相机陷阱会遇到部署时未知的物种。我们提出了Scout,一个自主的开放世界识别系统,它间歇性地调用云端VLM来向紧凑的边缘模型教授新类别。仅给定部署位置和空的站点帧,Scout自主地将VLM识别的每个物种转化为持久的、站点条件下的识别能力,并集成到资源高效的边缘模型中,无需预定义物种列表、人工标注或手动调优。在三个区域的30个相机陷阱部署中,使用NVIDIA Jetson Orin Nano,Scout的准确率保持在给定预定义物种列表的模型的0.1%-2.5%以内。对于初始类别集合之外的物种,Scout实现了53.7%-59.1%的准确率,而完全云端卸载为56.5%-65.1%,同时部署能耗降低了59%-71%。
英文摘要
Large vision-language models (VLMs) enable recognition beyond a fixed class set, but their computational demands prevent them from running on many edge devices. Cloud offload makes this capability accessible, but sending every image consumes scarce bandwidth and communication energy. We ask how to bring the open-world recognition capability of VLMs to the edge while operating within tight compute, energy, and bandwidth budgets. Wildlife monitoring provides a natural setting for exploring this question because camera traps encounter species not known at deployment. We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model. Given only the deployment location and empty site frames, Scout autonomously turns each species identified by the VLM into persistent, site-conditioned recognition capability in a resource-efficient edge model, without a predefined species list, human labeling, or manual tuning. Across 30 camera-trap deployments in three regions on an NVIDIA Jetson Orin Nano, the accuracy of Scout remains within 0.1-2.5% of a model given a predefined species list. On species outside its initial class set, Scout achieves 53.7-59.1% accuracy, compared with 56.5-65.1% for full cloud offload, while using 59-71% less deployment energy.
CommentsUnder Review