SSP:面向自动驾驶VLA模型的事件匹配Syn2Sim2Phy跨域评估框架
SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models
- College of Automotive and Energy Engineering, Tongji University(同济大学汽车与能源工程学院)
- Tongji Automotive Design & Research Institute Co., Ltd.(同济汽车设计研究院有限公司)
- Hubei Jingchu Humanoid Robot Co., Ltd.(湖北荆楚人形机器人有限公司)
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出SSP框架,通过锚定同一安全关键交互事件实现自动驾驶VLA模型的跨域评估,在合成、仿真、真实域中测得不同模型的性能差异,为VLA行为提供可复现的评估方法。
中文摘要 AI 辅助
用于自动驾驶的视觉-语言-动作(VLA)模型可协同完成场景解读、基于语言的推理及驾驶轨迹生成。现有评估常采用独立选取的合成、仿真及真实场景数据,因此测得的性能差距可能受场景内容变化而非真正的域敏感性混淆。本文提出SSP(Synthetic-Simulation-Physical,合成-仿真-真实),一种事件匹配的Syn2Sim2Phy评估框架,将跨域比较锚定到同一安全关键交互事件。从合成长尾视频出发,SSP构建经验证的事件规范,保留道路拓扑、参与角色、相对运动、冲突演化、通行顺序、响应约束及事件阶段。随后在CARLA平台及封闭试验场构建特定平台的实现,仅在传输审计确认强制事件属性保留后进行评估。SSP将OpenEMMA、LLaViDA及Alpamayo-R1的异构输出映射到通用语义槽和1秒轨迹窗口,以评估输出有效性、语义准确性、关键交互识别、轨迹质量及风险响应。在切入(Cut-in)和弱势道路使用者横穿场景中,合成、仿真、真实域的宏平均VLA综合能力得分分别为0.259、0.291、0.325,最佳域随场景变化;Alpamayo-R1、OpenEMMA、LLaViDA的得分分别为0.405、0.338、0.131。SSP提供可复现的场景传输链及基于证据的VLA行为评估,不假设真实域普遍更优。
英文摘要
Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.