发表机构
Key Laboratory of Water Big Data Technology of Ministry of Water Resources, Hohai University; College of Computer Science and Software Engineering, Hohai University; School of Cyber Science and Technology, Sun Yat-sen University; School of Artificial Intelligence, Shenzhen Technology University(河海大学水资源大数据技术教育部重点实验室; 河海大学计算机科学与软件工程学院; 中山大学网络空间安全学院; 深圳职业技术大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SLaFE框架,利用大语言模型将交通法规转为结构化约束,并通过协同进化模糊测试优化场景,在Apollo平台上触发全部10种违规,优于现有方法。
AI 中文摘要
以成本效益高且高效的方式确保自动驾驶系统(ADS)的安全性仍然是一个关键挑战。现有的基于法规引导的场景生成方法通常局限于狭窄的法律规则子集,导致场景多样性不足,而基于搜索的方法往往难以处理庞大且稀疏的搜索空间。为解决这些局限性,我们提出了SLaFE(基于协同进化的语义法规建模与模糊测试),一种新颖的交通违规场景生成框架,旨在系统性地评估ADS的安全性。SLaFE利用大语言模型(LLM)的推理能力将交通法规转化为结构化的场景约束。随后,通过一种协同进化模糊测试算法对这些场景进行优化,该算法探索参数空间以识别可能触发ADS异常行为的边界情况。我们在LGSVL模拟器中的Apollo平台上,使用十项真实世界的交通法规对SLaFE进行了评估。实验结果表明,SLaFE成功触发了所有10种类型的交通法规违规(10/10),优于现有最佳方法VioHawk(9/10),而其他方法检测到的违规类型不超过3种。此外,SLaFE实现了每种法规类型平均5.1分钟的触发时间,显著快于VioHawk(9.0分钟)及其他基线方法。这些结果凸显了SLaFE在发现用于ADS测试的多样且关键的违规场景方面的有效性。
英文摘要
Ensuring the safety of autonomous driving systems (ADS) in a cost-effective and efficient manner remains a critical challenge. Existing law-guided scenario generation approaches are typically limited to a narrow subset of legal rules, resulting in insufficient scenario diversity, and search-based methods often struggle with large and sparse search spaces. To address these limitations, we propose SLaFE (Semantic Law Modeling and Fuzzing based on Cooperative Evolution), a novel traffic violation scenario generation framework designed to systematically evaluate the safety of ADS. SLaFE harnesses the reasoning capabilities of large language models (LLMs) to convert traffic laws into structured scenario constraints. These scenarios are then optimized via a cooperative evolutionary fuzzing algorithm that explores the parameter space to identify boundary cases likely to trigger abnormal ADS behaviors. We evaluate SLaFE on the Apollo platform within the LGSVL simulator using ten real-world traffic regulations. Experimental results show that SLaFE successfully triggered all 10 types of traffic law violations (10/10), outperforming the best existing method, VioHawk (9/10), while others detected no more than 3. Moreover, SLaFE achieved an average triggering time of 5.1 minutes per law type, significantly faster than VioHawk (9.0 minutes) and other baselines. These results highlight SLaFE's effectiveness in discovering diverse and critical law-violating scenarios for ADS testing.