SAFARI:面向LLM辅助危害分析与风险评估的工业基准
SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment
浏览论文内容
中文总结 AI 辅助
提出首个工业基准SAFARI,含3000个ISO 26262汽车HARA案例,评估LLM在危害分析与风险评估中的表现,发现模型在风险分类上较弱(最佳ASIL宏F1仅0.261),并定位主要失败原因,为专家监督提供指引。
中文摘要 AI 辅助
大型语言模型(LLM)在安全关键工程中的应用日益受到关注,但它们在受监管的功能安全流程中的可靠性仍未得到充分探索。我们提出了SAFARI(Safety-Aware Functional Automotive Risk Inference,安全感知的功能汽车风险推断),这是首个面向ISO 26262标准下LLM辅助汽车危害分析与风险评估(HARA)的工业基准。该基准包含3,000个去标识化的工业HARA案例,并评估两个耦合任务:开放式危害分析和基于标准的风险评估。为了评估开放式HARA产物,我们提出了首个参考锚定的LLM-as-a-judge协议,该协议与专家评估具有高度相关性。对九个前沿LLM的实验表明,模型通常能生成看似合理的危害描述,但在ISO 26262风险分类方面表现薄弱,最佳ASIL宏F1分数仅为0.261。链式思维提示带来的收益有限,且常常降低分类风险评估的性能。错误分析进一步将主要失败定位到危害生成过程中场景关键上下文遗漏以及风险评估中的可控性误判,这表明专家监督应集中于此。数据集可通过此https URL获取。
英文摘要
Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from https://github.com/xixi47520-hash/HARA.