发表机构
EPFL; Detectium(洛桑联邦理工学院; Detectium)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过师生知识蒸馏框架压缩领域专用视觉语言模型,实现设备端火灾检测,实验表明Qwen2.5-0.5B在准确性、延迟和内存间取得最佳平衡,为资源受限场景提供部署指导。
AI 中文摘要
视觉语言模型(VLM)通过推理场景的语义上下文来减少误报,为传统火灾检测系统提供了一种有前景的替代方案,但其庞大的模型规模使得在嵌入式火灾传感器上的部署不切实际。在本文中,我们研究了如何压缩领域专用的视觉语言模型,以实现完全设备端部署,同时不丢失火灾检测所需的安全关键行为。我们开发了一个师生知识蒸馏框架,其中针对火灾理解进行微调的大型视觉语言模型可以被蒸馏为轻量级学生模型。跨多个视觉语言模型家族和模型规模的实验表明,紧凑的学生模型保留了其教师模型的大部分火灾理解能力。我们进一步将蒸馏后的模型部署在我们的商用Detectium火灾检测传感器上,并联合评估推理准确性、延迟和内存使用情况。结果表明,压缩和部署不仅影响准确性,还影响模型故障模式,其中Qwen2.5-0.5B提供了最强的整体部署权衡。我们的研究结果为在资源受限、安全关键的场景中部署领域专用的视觉语言模型提供了更广泛的指导。
英文摘要
Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers' fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.