Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
跨模态上下文感知学习:用于遥感中视觉提示引导的多模态图像理解
机构 * College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学) ; School of Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人学院及机器人视觉感知与控制技术国家工程研究中心,湖南大学) ; School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算与人工智能学院,西南交通大学)
专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV
AI总结 CLV-Net通过跨模态上下文感知学习,利用视觉提示引导遥感多模态图像理解,提升目标识别精度和用户意图对齐能力。
Comments 12 pages, 5 figures