发表机构
Beihang University; The Chinese University of Hong Kong (Shenzhen); The Chinese University of Hong Kong; The Hong Kong University of Science and Technology (Guangzhou); Shanghai Artificial Intelligence Laboratory(北京航空航天大学; 香港中文大学(深圳); 香港中文大学; 香港科技大学(广州); 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VLALight,一种轻量级端到端视觉-语言-动作框架,直接映射交叉口观测到信号动作,以0.5B参数实现实时紧急感知交通信号控制,降低紧急等待时间21.1%。
AI 中文摘要
交通信号控制(TSC)对于缓解城市拥堵至关重要。近期视觉-语言模型(VLM)的进展使得对交叉口场景的更丰富解读成为可能,为视觉上下文感知的TSC开辟了新机遇。然而,模块间的松散耦合和重复信息转换可能导致细粒度视觉细节的丢失,而顺序推理则引入了显著延迟。为解决这些局限,我们提出VLALight,一种轻量级端到端视觉-语言-动作框架,直接将交叉口观测和信号相位信息映射到离散信号动作。为处理TSC的多视角特性,VLALight将多个方向摄像头视图组合为统一视觉输入,并使用文本指令建立其与交通流向和信号相位的对应关系。该设计使得仅用0.5B参数的紧凑模型即可实现直接动作预测,无需中间图像到文本描述或手工制作的交通状态表示。实验表明,VLALight在所有对比方法中提供了最佳的紧急车辆服务,将级联VLMLight的汇总紧急等待时间降低了21.1%,同时在本地硬件上实时运行,并泛化到未见过的交叉口拓扑和交通流模式。
英文摘要
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.
Comments9 pages, 7 figures