arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLALight:面向交通信号控制的视觉-语言-动作模型

VLALight: A Vision-Language-Action Model for Traffic Signal Control

Pan Zhang, Siqi Lai, Kemu Dong, Hao Liu

arXiv 2609.36934首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Dalian University of Technology(香港科技大学(广州); 大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VLALight是首个端到端视觉-语言-动作交通信号控制模型,通过多视角视频直接映射信号动作,采用两阶段冷启动训练和协同强化学习,在七个真实数据集上超越现有基线,实现高效自适应推理。

AI 中文摘要

交通信号控制(TSC)对于改善城市交通流动性和减少拥堵至关重要。尽管路侧摄像头在信号交叉口广泛部署,并提供不断变化的交通的丰富视觉观测,但现有的TSC方法通常依赖人工设计的交通状态或独立的感知模块,在物理观测与控制决策之间造成了差距。我们提出了VLALight,这是首个用于从多视角路侧视频进行端到端交通信号控制的视觉-语言-动作(VLA)模型。VLALight通过多目标时空交通推理和跨交叉口的拓扑感知协同感知,直接将视觉观测映射为协调的信号动作。为了建立这一能力,我们开发了一种两阶段监督冷启动训练策略,用于视觉交通理解和信号决策,随后进行协同智能体强化学习,联合优化局部控制和网络级交通效率。此外,VLALight引入了自适应快速和慢速推理模式,使策略仅在额外思考能提供足够控制收益时才分配更深层次的推理。通过平衡的模态感知轨迹展开和相对优势优化,VLALight学会了权衡决策质量和推理成本。在三个城市网络的七个真实交通流数据集上进行的大量实验表明,VLALight持续优于基于交通工程、基于强化学习(RL)和基于LLM/VLM的基线方法。消融研究验证了协同感知、网络级优化和自适应推理的有效性。这些结果展示了VLA模型在现实物理交通控制中的潜力。我们的项目可在以下网址获取:https://this URL。

英文摘要

Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑