机构
*
Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
;
School of Biomedical Engineering, Southern Medical University(南方医科大学生物医学工程学院)
;
Singapore University of Technology and Design(新加坡科技与设计大学)
;
University of Auckland(奥克兰大学)
;
University of Science and Technology of China(中国科学技术大学)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
College of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
在可观察之前看见:通过视觉语言模型在自动驾驶中的潜在风险推理
Jiaxin Liu, Xiangyu Yan, Liang Peng, Lei Yang, Lingjun Zhang, Yuechen Luo, Yueming Tao, Ashton Yu Xuan Tan, Mu Li, Lei Zhang, Ziqi Zhan, Sai Guo, Hong Wang, Jun Li
机构
*
School of Vehicle and Mobility, Tsinghua University(车辆与移动系统学院,清华大学)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
Video-CoM:通过操作链进行交互式视频推理
Hanoona Rasheed, Mohammed Zumri, Muhammad Maaz, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
机构
*
Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
;
University of California Merced(加州梅尔 Ced 大学)
;
Google Research(谷歌研究院)
;
Linköping University(林奈大学)
;
Australian National University(澳大利亚国立大学)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
像人类一样适应:具有测试时推理的元认知代理
Yang Li, Zhiyuan He, Yuxuan Huang, Zhuhanling Xiao, Chao Yu, Meng Fang, Kun Shao, Jun Wang
机构
*
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
University of Oxford(牛津大学)
;
Tsinghua University(清华大学)
;
University of Liverpool(利物浦大学)
;
University College London(伦敦大学学院)
SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing
SAMChat:引入链式推理和GRPO以增强小规模遥感遥感小语言模型
Aybora Koksal, A. Aydin Alatan
机构
*
Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中东部技术大学(METU)电子与电气工程系)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
CommentsAccepted to Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) Special Issue on Foundation and Large Vision Models for Remote Sensing. Code and dataset are available at https://github.com/aybora/SAMChat
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
;
Sofia University(索菲亚大学)
;
University of Macau(澳门大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中
视觉推理
:vision language model(abstract);VLM(abstract);分类 cs.CV
机构
*
School of Computer Science and Technology, Anhui University, Hefei 230601, China(安徽大学计算机科学与技术学院)
;
School of Artificial Intelligence, Anhui University, Hefei 230601, China(安徽大学人工智能学院)
;
Department of Computer Science and Information Technology, La Trobe University, Bendigo, Australia(拉筹伯大学计算机科学与信息技术系)