机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Institute of AI for Industries, Chinese Academy of Sciences(中国科学院工业人工智能研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉定位与Grounding
:visual reasoning(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
从词语到波长:用于少样本多光谱目标检测的视觉语言模型
Manuel Nkegoum, Minh-Tan Pham, Élisa Fromont, Bruno Avignon, Sébastien Lefèvre
机构
*
Univ Bretagne Sud, IRISA, UMR 6074(布列塔尼大学,IRISA,UMR 6074)
;
Univ Rennes, IRISA, UMR 6074(里尔大学,IRISA,UMR 6074)
;
ATERMES
;
UiT The Arctic University of Norway(北欧大学)
CommentsarXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理
Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun
机构
*
State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)
专题命中
视觉定位与Grounding
:visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery
Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki
机构
*
Department of Architecture, National University of Singapore(建筑系,新加坡国立大学)
;
Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学)
;
School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学)
;
Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)
专题命中
视觉定位与Grounding
:vision-language model(abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV
Journal refISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025
Spatial Preference Rewarding for MLLMs Spatial Understanding
Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang, Ling Shao, Shijian Lu
机构
*
S-Lab, Nanyang Technological University(南洋理工大学S实验室)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Sensetime Research(商汤科技研究院)
;
Zhejiang University of Technology(浙江工业大学)
;
UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构
*
South China University of Technology(华南理工大学)
;
Institute for Infocomm Research (I 2 R), A*STAR(信息通信研究所(I 2 R),A*STAR)
;
WeChat Vision, Tencent Inc.(微信视觉,腾讯公司)
;
Foshan University(佛山大学)
;
Nanyang Technological University(南洋理工大学)
;
National University of Singapore(新加坡国立大学)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
Changde Du, Yizhuo Lu, Zhongyu Huang, Yi Sun, Zisen Zhou, Shaozheng Qin, Huiguang He
机构
*
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术重点实验室,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学)
;
State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(认知神经科学与学习国家重点实验室,北京师范大学)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构
*
Hong Kong Baptist University(香港 Baptist 大学)
;
Sichuan University(四川大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
Qihang Fan, Huaibo Huang, Mingrui Chen, Ran He
机构
*
MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(自动化研究所,中国科学院,北京)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京)
专题命中
视觉定位与Grounding
:LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong, Zechao Li
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computing, Australian National University(澳大利亚国立大学计算学院)
;
Institute for Advanced Study in Mathematics, Harbin Institute of Technology(哈尔滨工业大学数学高等研究院)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D Watson, Levente J Klein, Fahad Shahbaz Khan, Salman Khan
机构
*
IBM Research(IBM研究院)
;
Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
Ruizhe Ou, Yuan Hu, Fan Zhang, Jiaxin Chen, Yu Liu
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Peking University(北京大学)
;
Peking University Ordos Research Institute of Energy(北京大学鄂尔多斯能源研究院)