SAFIRE:多模态大语言模型中细粒度火灾与烟雾理解的安全关键基准
SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs
浏览论文内容
中文总结 AI 辅助
针对多模态大语言模型在火灾烟雾理解中的安全关键评估空白,提出含83K图像、193K问答的SAFIRE基准,并验证数据微调可大幅提升分类性能。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在视觉-语言任务上展现出强劲进展,但其在安全关键场景中的可靠性仍未得到充分探索。火灾与烟雾理解对公共安全和灾害响应至关重要,但现有大多数基准缺乏多样化的真实世界场景和上下文感知评估。我们提出了SAFIRE,一个用于MLLMs火灾与烟雾理解的大规模基准,包含来自20个场景的83K张带标题图像,以及从9.7K图像子集生成的193K个多项选择视觉问答(MCVQA),涵盖从基础感知到高阶推理的10个评估维度。一个由GPT-5.4辅助的多阶段验证流程结合MLLM多数投票确保了标注质量。评估十个开源MLLMs(8B-38B)得出平均准确率为61.9%,暴露了安全关键推理方面的重大差距。我们进一步表明,仅使用我们领域特定数据的7%来适配视觉编码器,即可将火灾场景分类准确率从20.1%提升至64.5%,这表明即使数据量有限,精心策划的数据也能带来显著收益。所有数据集、模型和代码均可在该https URL获取。
英文摘要
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited. All datasets, models, and code are available at https://risys-lab.github.io/SAFIRE/.