TinyRS-R1: Compact Multimodal Language Model for Remote Sensing
TinyRS-R1:用于遥感的紧凑多模态语言模型
Aybora Koksal, A. Aydin Alatan
机构
*
Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中欧技术大学(METU)电子与电气工程系)
CommentsAccepted to IEEE Geoscience and Remote Sensing Letters (GRSL). Code, models, and the captions for datasets are available at https://github.com/aybora/TinyRS
LLM-Guided Indoor Navigation with Multimodal Map Understanding
Alberto Coffrini, Paolo Barsocchi, Francesco Furfari, Antonino Crivello, Alessio Ferrari
机构
*
Institute of Information Science and Technologies (ISTI)(信息科学与技术研究所)
;
National Research Council of Italy (CNR)(意大利国家研究理事会)
;
Department of Computer Science(计算机科学系)
;
University of Pisa(比萨大学)
;
University College Dublin (UCD)(都柏林大学学院)
;
School of Computer Science(计算机科学学院)
机构
*
Kyoto University(京都大学)
;
NII LLMC(日本国立信息与通信技术研究所语言模型中心)
;
RIKEN AIP(日本理化学研究所先进理工研究所)
;
Case Western Reserve University(凯斯西储大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Osaka(大阪大学)
;
University of Tokyo(东京大学)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
LatentPilot: 通过潜意识视觉推理进行场景感知的视觉-语言导航
Haihong Hao, Lei Chen, Mingfei Han, Changlin Li, Dong An, Yuqiang Yang, Zhihui Li, Xiaojun Chang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Stanford University(斯坦福大学)
;
Amap, Alibaba Group(阿里巴巴集团高德地图)
;
Shanghai AI Laboratory(上海人工智能实验室)