FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
机构 * Xi’an Jiaotong University(西安交通大学) ; Amap, Alibaba Group(阿里巴巴集团) ; DAMO Academy, Alibaba Group(阿里巴巴达摩院)
专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025 as Spotlight Presentation. Code: https://github.com/MIV-XJTU/FSDrive