面向安全关键驾驶中威胁感知控制的语言结构化关系Q学习
Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving
浏览论文内容
中文总结 AI 辅助
该研究针对安全关键驾驶,提出语言结构化关系Q学习(ERQ-Net),经2500个场景测试,提升了威胁感知能力但存在识别-控制差距,为相关策略学习提供了新视角。
中文摘要 AI 辅助
基于自然语言的场景生成为描述罕见且复杂的驾驶交互提供了直观手段,但目前仍不确定使用语言结构化数据训练是否能生成真正自适应的控制策略。我们提出语言结构化关系Q学习,通过以自我为中心的关系Q网络(ERQ-Net)实现,该网络可从动态交通图中联合学习车辆间相关性与动作价值。训练期间,语言描述定义周围车辆的行为,而提示和语义角色对策略隐藏,因此ERQ-Net必须仅从可观测的运动学数据和交互中推断威胁相关性。在2500个安全关键场景中,语言结构化训练将测试成功率从49%-52%提升至55%-58%,并将针对对手的关注度从1.2倍提高到2.1倍,展现出涌现的威胁感知能力。然而,这种表征增益并未始终转化为自适应控制:训练后的策略表现与最佳恒定动作相似,而一组简单策略可解决76%的场景。我们将这种差异形式化为“识别-控制差距”,并表明奖励重加权和边际塑造无法消除由此产生的策略崩溃。对真实度、关键度、语义准确性以及状态-接口表征向CARLA的迁移的评估,进一步凸显了安全关键驾驶场景中语言结构化关系策略学习的优势与局限。
英文摘要
Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which jointly learns inter-vehicle relevance and action values from dynamic traffic graphs. Language descriptions define surrounding-vehicle behaviours during training, while prompts and semantic actor roles are hidden from the policy. ERQ-Net must therefore infer threat relevance solely from observable kinematics and interactions. Across 2,500 safety-critical scenarios, language-structured training improves test success from 49-52% to 55-58% and increases adversary-focused attention from 1.2x to 2.1x, demonstrating emergent threat awareness. However, this representational gain does not consistently translate into adaptive control: trained policies perform similarly to the best constant action, while a portfolio of simple policies solves 76% of scenarios. We formalise this discrepancy as a recognition-control gap and show that reward reweighting and margin shaping do not eliminate the resulting policy collapse. Evaluations of realism, criticality, semantic accuracy, and transfer of state-interface representations to CARLA further highlight both the strengths and the constraints of language-structured relational policy learning in safety-critical driving scenarios.
发表机构
- Edge Hill University(埃奇希尔大学)
机构由 AI 辅助整理,请以论文原文为准。