基于大语言模型的汽车软件验证传感器级故障诊断
Sensor-Level Fault Diagnosis for Automotive Software Validation Using Large Language Models
浏览论文内容
中文总结 AI 辅助
该研究提出两阶段框架,用经4-bit低秩适配的LLM对HIL平台的汽车传感器记录做故障诊断,最小模型准确率达81.6%,适配循环可在单商用加速器运行。
中文摘要 AI 辅助
在硬件在环(HIL)平台上进行的汽车软件预系列验证会产生大量多变量传感器记录,针对功能安全要求的评估在大规模试验中超出了人工审查的承受能力。基于阈值的工具仅报告发生了偏差,但既无法识别其性质也无法定位其来源;而数据驱动分类器虽然准确,但依赖大量带标注数据集,且返回的决策不透明,不符合ISO 26262要求的可追溯性。本研究探究开源指令微调大语言模型(LLM)在给定传感器行为文本描述的情况下,是否能作为验证循环中数据高效且可解释的故障检测与诊断引擎。提出了两阶段框架:在dSPACE实时平台上进行自动化需求检查,首先分离出违反安全需求的记录,仅对这些记录进行检查;将信号的滑动窗口简化为统计、关系和上下文描述符,嵌入固定提示中,再由经4-bit低秩适配的LLM映射为故障位置及书面说明。对参数规模在20亿至80亿之间的四个模型系列进行适配,并在涵盖6类注入故障的汽油发动机案例研究中测试:最小模型与最大模型的准确率均达81.6%,而同等规模的另一模型未收敛,表明任务特定适配下的诊断能力取决于收敛性而非参数数量,且整个适配与评估循环可在单个商用加速器上运行。
英文摘要
The pre-series validation of automotive software on hardware-in-the-loop (HIL) platforms produces large volumes of multivariate sensor recordings whose assessment against functional safety requirements exceeds what manual review can sustain at campaign scale. Threshold-based tooling reports that a deviation has occurred but neither identifies its nature nor locates its source, while data-driven classifiers, although accurate, rely on large labelled datasets and return opaque decisions that sit uneasily with the traceability demanded by ISO 26262. This study examines whether open-source instruction-tuned large language models (LLMs), given a textual description of sensor behaviour, can serve as data-efficient and interpretable engines for fault detection and diagnosis inside the validation loop. A two-phase framework is proposed: automated requirement checking on a dSPACE real-time platform first isolates the recordings that violate a safety requirement, and only these are inspected, with sliding windows of the signals reduced to statistical, relational, and contextual descriptors, embedded in a fixed prompt, and mapped by a 4-bit low-rank-adapted LLM to a fault location accompanied by a written justification. Four model families ranging from two to eight billion parameters were adapted and tested on a gasoline-engine case study spanning six injected fault classes. The smallest model matched the largest at 81.6\% accuracy, whereas a comparably sized model failed to converge, indicating that diagnostic competence under task-specific adaptation follows convergence rather than parameter count, with the entire adapt-and-evaluate cycle fitting on a single commodity accelerator.