arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Theia:用于无数据蒸馏的Incidents1M数据集的大规模多模态字幕生成与自动验证

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

Simone Giano, Lorenzo Severini, Alessandro Galdelli, Adriano Mancini

arXiv 2607.28269首次发表:更新:

发表机构

Università Politecnica delle Marche(马尔凯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对灾害领域多模态数据集的缺陷,提出方法构建并自动验证Incidents1M数据集,生成高保真字幕,其语义一致性达78.65/100,为跨模态知识蒸馏提供了可扩展框架。

AI 中文摘要

视觉语言模型(VLMs)在灾害管理等关键领域的部署需要高质量多模态数据集,尤其是通过无数据知识蒸馏(DFKD)进行知识迁移时。然而,该领域现有数据集要么完全缺乏描述性文本(如Incidents1M),要么存在严重的图文语义错位(如CrisisMMD)。本研究提出一种构建并自动验证灾害应对大规模多模态数据集的新方法:从仅含视觉内容的Incidents1M出发,成功恢复10万张图像,并采用两种不同的Qwen3.5架构生成高保真文本描述——4B稠密模型与35B混合专家(MoE)模型;为确保生成的字幕能为DFKD提供可靠语义锚点,引入利用Qwen3.5-9B的盲图像LLM评估流水线,通过故意向评估器隐藏原始图像,准确模拟无数据蒸馏期间学生模型的模态差距。对173179个标签对的评估显示,两种架构间语义一致性达78.65/100;此外,自动评估揭示了保守的字幕行为,表现为高准确率(77.6%)与低召回率(46.0%),这既最小化了假阳性噪声,又暴露了原始真值中潜在的人工标注不一致性。本研究提供了可扩展的LLM验证多模态数据集与可复现框架,以推进跨模态知识蒸馏。

英文摘要

The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Data-Free Knowledge Distillation (DFKD). However, existing datasets in this domain either entirely lack descriptive text, such as Incidents1M, or suffer from severe text-image semantic misalignment, such as CrisisMMD. In this work, we present a novel methodology to construct and automatically validate a large-scale multimodal dataset for disaster response. Starting from the vision-only Incidents1M, we successfully recovered 100,000 images and generated high-fidelity textual descriptions using two distinct Qwen3.5 architectures: a 4B dense model and a 35B Mixture-of-Experts (MoE) model. To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline leveraging Qwen3.5-9B. By intentionally obscuring the original image from the judge, this evaluator accurately simulates the modality gap of the student model during data-free distillation. Our evaluation across 173,179 label pairs demonstrates a high semantic agreement (78.65/100) between the two architectures. Furthermore, the automated evaluation reveals a conservative captioning behaviour, characterized by a high Precision (77.6%) and low Recall (46.0%). This minimizes the false positive noise, while simultaneously exposing underlying human annotation inconsistencies in the original ground truth. This work provides a scalable, LLM-validated multimodal dataset and a reproducible framework to advance cross-modal knowledge distillation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑