发表机构
FAU Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对VLA驾驶模型无法在推理时重定向注意力的问题,在Qwen3-VL主干上添加预softmax注意力偏差,实验验证了其对轨迹的引导效果及层相关特性。
AI 中文摘要
视觉-语言-动作(VLA)驾驶模型将推理阶段与基于扩散的轨迹解码器耦合,但在推理时无法在不重新训练的情况下直接将注意力重定向到安全关键的交通参与者。我们在Alpamayo-R1的Qwen3-VL主干网络中,对检测器定位的交通参与者的视觉标记应用了有界加法预softmax注意力偏差,该偏差作为无权重更新的故障开放前向预钩子使用。在Physical AI World Model合成数据集的50个变道场景中,轨迹解码器在偏差幅度上呈现单调剂量响应,且在每个测试幅度下均与配对的零偏差对照组分离。其平均位移约为17厘米,钳位时横向偏移高达约140厘米。层消融实验表明,动作相关信号位于后期层,效果随钩子层数增加而提升(前8层为2.0厘米,全部36层为67.6厘米)。逐调用注入审计解释了因果链文本从未改变的原因:基于掩码的偏差在该服务栈中从未到达推理通路,因此不变性是暴露验证而非鲁棒性验证。引导后的轨迹倾向于向被关注的参与者偏移,表明该偏差控制模型的关注位置而非编码目标行为。
英文摘要
Vision-language-action (VLA) driving models couple a reasoning stage with a diffusion-based trajectory decoder, but do not give a direct way to redirect attention toward safety-critical actors at inference time without retraining. We studied a bounded additive pre-softmax attention bias on the visual tokens of detector localized traffic actors on Alpamayo-R1's Qwen3-VL backbone. It is applied as a fail open forward pre-hook with no weight changes. On 50 lane-change scenarios from the Physical AI World Model Synthetic dataset. The trajectory decoder shows a monotonic dose response in the bias magnitude, separate from a paired zero bias control at every tested magnitude. It reaches $\approx 17$\,cm mean displacement with lateral shifts up to $\sim 140$\ cm at the clamp. A layer ablation places the action-relevant signal in late layers, where the effect increases with the number of hooked layers (2.0cm for the first 8 layers; 67.6cm for all 36). A per call injection audit explains why the Chain-of-Causation text never changes. The mask based bias never reaches the reasoning pathway in this serving stack, so the invariance is verified exposure, not robustness. Steered trajectories tend to shift toward the attended actor, suggesting the bias governs where the model looks rather than encoding a target behavior.
CommentsAttention Steering, Vision-Language-Action, AutonomousDriving, Inference-Time Intervention
Journal refEuropean Conference on Computer Vision 2026