发表机构
University of Michigan-Dearborn(密歇根大学迪尔伯恩分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Teach-to-Crash闭环双LLM框架,通过教师指导搜索和学生生成场景,在CARLA中实现高碰撞发现率与多样性,提升ADS安全测试效率。
AI 中文摘要
在仿真环境中验证自动驾驶系统(ADS)需要能够发现罕见、安全关键的故障的测试架构,同时生成可执行、多样化且对下游故障分析有用的场景。我们提出了Teach-to-Crash,一个闭环测试框架,它结合了受限的以自我为中心的场景表示、停滞感知的搜索控制以及用于自适应故障发现的双大语言模型(LLM)架构。一个高推理能力的教师LLM充当自适应搜索控制器,而一个低推理能力的学生LLM以严格的JSON模式生成模拟器可执行的场景。教师仅在滚动碰撞率和碰撞时间指标停滞时进行干预,提供战略指导以重新引导搜索。在CARLA案例研究中,使用两种实验设置来改变自我车辆的速度策略,Teach-to-Crash实现了最高的碰撞命中率(90.79%)、最短的平均碰撞时间(18.31秒)以及具有竞争力的碰撞发现率(136.21)。PAFOT获得了更高的平均CDR(179.44),但方差显著更大。Teach-to-Crash还产生了最高的多样性(0.547),并且在CARLA交通管理器控制器上的两种设置中平均,提供了最高的基于可避免性的有用性代理(60.04%)。这些结果在评估的CARLA范围内提供了证据,表明闭环双LLM推理可以在受限的可执行程序空间上引导基于对抗仿真的测试,生成频繁、结构多样且被评估为更常可避免的故障。
英文摘要
Validating Autonomous Driving Systems (ADS) in simulation requires testing architectures that can discover rare, safety-critical failures while generating scenarios that are executable, diverse, and useful for downstream failure analysis. We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery. A high-reasoning Teacher LLM acts as an adaptive search controller, while a low-reasoning Student LLM emits simulator-executable scenarios in a strict JSON schema. The Teacher intervenes only when rolling collision rate and time-to-collision metrics stagnate, providing strategic guidance to redirect the search. In a CARLA case study with two experimental setups that vary the ego vehicle's speed policy, Teach-to-Crash achieves the highest Collision Hit Rate (90.79%), the shortest mean Time-to-Collision (18.31 s), and a competitive Collision Discovery Rate (136.21). PAFOT attains a higher mean CDR (179.44), but with substantially larger variance. Teach-to-Crash also yields the highest diversity (0.547) and, averaged across both setups on the CARLA Traffic Manager controller, the highest avoidability-based usefulness proxy (60.04%) among the compared methods. These results, within the evaluated CARLA scope, provide evidence that closed-loop dual-LLM reasoning can steer adversarial simulation-based testing over a constrained executable program space, generating failures that are frequent, structurally diverse, and assessed as more frequently avoidable.
Comments41st IEEE/ACM International Conference on Automated Software Engineering (ASE) AgenticDev (2026)