AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) ; CASIA(中国科学院自动化所) ; Alibaba Cloud(阿里云) ; School of Intelligence Science and Technology(智能科学与技术学院) ; CAIR(中国科学院香港创新研究院)
专题命中 视觉定位与Grounding :vision-language model(title);visual language model(abstract);分类 cs.CV、cs.AI