在多传感器物理危害评估中对大语言模型进行基准测试
Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
浏览论文内容
中文总结 AI 辅助
研究五个大语言模型对多传感器物理危害数据的评估,通过1800次API调用测试60个场景,发现模型在多传感器场景表现不佳,单传感器场景表现良好,且结构化表格格式优势不明显,为从业者提供了模型应用参考。
中文摘要 AI 辅助
我们进行了一项实证基准测试,以评估五个大语言模型如何评估多传感器物理危害数据。在温度0.0下进行1800次API调用,测试了跨多传感器联合评估、响应比例性和模式消歧三类的60个场景。我们发现,在多个传感器同时低于其各自安全极限的测试场景中,所有测试模型均未持续产生预防性警告信号,而在单传感器阈值违规方面达到了近乎完美的准确率。所有五个模型在A类多传感器场景中得分接近零,而在单传感器场景中表现强劲。结构化表格格式与普通散文相比没有一致优势;ChatGPT-4o在散文形式下表现明显更好。这些发现对在物理安全监测系统中部署测试模型的从业者有直接影响。
英文摘要
We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three categories - multi-sensor joint assessment, response proportionality, and pattern disambiguation - with 1,800 API calls at temperature 0.0, we find that all tested models consistently produced no precautionary warning signal across the tested scenarios where multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor threshold violations. All five models (ChatGPT-4o, Gemini 2.5 Flash, DeepSeek, Kimi, Llama 3.1 8B) score near zero on Category A multi-sensor scenarios (Q2: 0.000-0.208; Q3: 0.000-0.592) compared to strong performance on single-sensor scenarios (Category B Q1: 0.975-1.000). Structured tabular formatting shows no consistent advantage over plain prose; ChatGPT-4o performs significantly better under prose (p = 0.001). These findings have direct implications for practitioners deploying the tested models in physical safety monitoring systems.
发表机构
- Lovely Professional University(可爱专业大学)
机构由 AI 辅助整理,请以论文原文为准。