Towards Multimodal Social Conversations with Robots: Using Vision-Language Models
Ruben Janssens, Tony Belpaeme
机构
*
Ghent University–imec(根特大学–imec)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CL
CommentsAccepted at the workshop "Human - Foundation Models Interaction: A Focus On Multimodal Information" (FoMo-HRI) at IEEE RO-MAN 2025 (Camera-ready version)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
Enming Zhang, Bingke Zhu, Yingying Chen, Qinghai Miao, Ming Tang, Jinqiao Wang
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
Wuhan AI Research(武汉人工智能研究所)
;
Peng Cheng Laboratory(鹏城实验室)
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及交叉应用关键实验室(东南大学),中华人民共和国教育部,中国)
;
School of Software, Northwestern Polytechnical University(西北工业大学软件学院)
;
State Key Laboratory for Novel Software Technology and School of Intelligence Science and Technology, Nanjing University(新型软件技术国家重点实验室和南京大学智能科学与技术学院)
;
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data
Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima
机构
*
Lawrence Berkeley National Laboratory(伯克利国家实验室)
;
University of California, Irvine(加州大学尔湾分校)
;
University of California, Berkeley(加州大学伯克利分校)
;
Covalent Metrology(协力计量)
专题命中
视频多模态
:multimodal(abstract);分类 cs.CV
CommentsThis paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
Haojie Zhang, Yixiong Liang, Hulin Kuang, Lihui Cen, Zhe Qu, Yigang Cen, Min Zeng, Shichao Kan
机构
*
School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)
;
School of Automation, Central South University(自动化学院,中南大学)
;
School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo
机构
*
MedVisAI Lab Department of ECE Northwestern University(MedVisAI实验室 电子工程系 西北大学)
;
Institute for Infocomm Research (I 2 R) A*STAR, Singapore(信息与通信研究所(I 2 R)A*STAR,新加坡)
;
MedVisAI Lab Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室 李科田医学院,南洋理工大学)