Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
解释未见的:多模态视觉-语言推理用于地下矿难中的情境感知
机构 * Missouri University of Science and Technology(密苏里科学与技术大学) ; Washington State University(华盛顿州立大学)
专题命中 推理评测 :reasoning(title)
AI总结 本文提出MDSE框架,通过多模态视觉-语言推理提升地下矿难情境感知能力,实现更准确的描述生成。