发表机构
Rensselaer Polytechnic Institute; IBM T. J. Watson Research Center(伦斯勒理工学院; IBM T. J. 沃森研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出风险自适应边云视觉推理架构,通过车载交通评估触发云端VLM推理,在自动驾驶中减少54.1%云端请求量,同时保持任务成功率并降低AEB触发次数,实现通信效率与性能的平衡。
AI 中文摘要
云端部署的视觉语言模型(VLMs)比小型车载模型具备更强的上下文推理能力,但频繁的视觉上传会增加通信开销,并给战术决策带来网络与推理延迟。本文提出一种风险自适应边云架构,其中车载交通评估模块决定何时请求云端推理。车载VLM与轻量检测器捕捉时间性交通状况及路径相关危险,用于保守的本地响应与选择性云端访问;云端模型提供战术建议,而验证、车辆控制及自动紧急制动(AEB)仍在本地执行。在CARLA实验中,本方法达到了周期性云端访问的任务成功率,同时减少54.1%的云端请求,且自动紧急制动(AEB)触发次数更少。在延迟道路施工的消融实验中,语义事件会在下一次计划审核前触发请求。在三种模拟网络配置下,该方法持续降低云端流量,尽管变道耗时比周期性访问更长。因此,在上述实验中,车载交通评估可作为选择性VLM推理的实用触发机制。
英文摘要
Cloud-hosted vision-language models (VLMs) offer greater contextual reasoning capabilities than smaller onboard models, but frequent visual uploads increase communication overhead and add network and inference latency to tactical decisions. We present a risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested. An onboard VLM and a lightweight detector capture temporal traffic conditions and path-relative hazards for conservative local response and selective cloud access. The cloud model provides tactical advice, while validation, vehicle control, and automatic emergency braking remain local. In CARLA experiments, our method matched the task success rate of periodic cloud access while reducing cloud requests by 54.1% and recording fewer automatic emergency braking (AEB) activations. In a delayed-roadwork ablation, semantic events triggered requests before the next scheduled audit. Across three emulated network profiles, the method continued to reduce cloud traffic, although lane changes took longer than with periodic access. Onboard traffic assessment therefore served as a practical trigger for selective VLM inference in these experiments.
Comments7 pages, 4 figures, 5 tables