机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Eastern Institute of Technology(东区技术研究所)
;
Great Bay University(大湾大学)
;
Northeast Forestry University(东北林业大学)
;
National University of Singapore(新加坡国立大学)
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
UniDrive-WM:面向自动驾驶的统一理解、规划和生成世界模型
Zhexiao Xiong, Xin Ye, Burhan Yaman, Sheng Cheng, Yiren Lu, Jingru Luo, Nathan Jacobs, Liu Ren
机构
*
Bosch Research North America & Bosch Center for Artificial Intelligence (BCAI)(博世北美研究院与博世人工智能中心)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Arizona State University(亚利桑那州立大学)
;
Case Western Reserve University(凯斯西储大学)
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
用于病例级病理学概要报告生成的简单令牌高效视觉语言模型
Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini
机构
*
Department of Computer Science and Software Engineering (CSSE), Concordia University, Montreal, Canada(计算机科学与软件工程系(CSSE),康科迪亚大学,蒙特利尔,加拿大)
;
Axe Cancer, Centre de recherche du CHUM, Université de Montréal, Montreal, Canada(Axe癌症,CHUM研究中心,蒙特利尔大学,蒙特利尔,加拿大)
;
Institut de recherche en immunologie et cancérologie (IRIC), Université de Montréal(免疫学与癌症研究所(IRIC),蒙特利尔大学)
;
Mila - Quebec AI Institute, Montreal, Canada(魁北克AI研究所(Mila),蒙特利尔,加拿大)
机构
*
Chinese University of Hong Kong(香港中文大学)
;
Westlake University(西湖大学)
;
Southern Medical University(南方医科大学)
;
Jiangnan University(江南大学)
;
Dalian University of Technology(大连理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Copenhagen(哥本哈根大学)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
SMART: 基于音频增强多模态大模型的镜头感知视频时刻检索
An Yu, Weiheng Lu, Jian Li, Zhenfei Zhang, Yunhang Shen, Felix X. -F. Ye, Ming-Ching Chang
机构
*
Department of Computer Science, University at Albany - SUNY(University at Albany - SUNY 计算机科学系)
;
School of Software & Microelectronics, Peking University(北京大学软件与微电子学院)
;
Nanjing University(南京大学)
;
Xiamen University(厦门大学)
;
Department of Mathematics and Statistics, University at Albany - SUNY(University at Albany - SUNY 数学与统计学系)
专题命中
VLM训练与架构
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Do VLMs Align Better with Humans than LLMs during Natural Reading?
VLMs 在自然阅读中可能不会全局性地增强与人类的对齐性优于 LLMs
Jinzhou Wu, Zhengwu Ma, Jixing Li, Baoping Tang, Zitong Lu
机构
*
Department of Mechanical and Vehicle Engineering, Chongqing University(重庆大学机械与车辆工程学院)
;
Department of Linguistics and Translation, City University of Hong Kong(香港城市大学语言学与翻译系)
;
McGovern Institute for Brain Research, Massachusetts Institute of Technology(麻省理工学院麦戈文脑科学研究所)
机构
*
Wuhan University(武汉大学)
;
University of Exeter(埃克塞特大学)
;
Chinese Academy of Sciences(中国科学院)
;
Xi'an University of Electronic Science and Technology(西安电子科技大学)
Laurens Samson, Nimrod Barazani, Sennay Ghebreab, Yuki M. Asano
机构
*
Socially-Intelligent Artificial Systems Group, University of Amsterdam(智能社会人工智能系统组,阿姆斯特丹大学)
;
University of Amsterdam(阿姆斯特丹大学)
;
Fundamental AI Lab, University of Technology Nuremberg(基础人工智能实验室,纽伦堡技术大学)
专题命中
VLM训练与架构
:visual language model(title,abstract);VLM(abstract_cn);分类 cs.CV
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
基于置信度蒸馏的门控关系对齐用于高效视觉语言模型
Yanlong Chen, Amirhossein Habibian, Luca Benini, Yawei Li
机构
*
Department of Information Technology(信息科技系)
;
Electrical Engineering, ETH Zurich, Zurich, Switzerland(电气工程,苏黎世联邦理工学院,苏黎世,瑞士)
;
Qualcomm AI Research, Amsterdam, the Netherlands(高通人工智能研究,阿姆斯特丹,荷兰)
;
Department of Electrical, Electronic and Information Engineering(电气、电子与信息工程系)
;
University of Bologna, Bologna, Italy(博洛尼亚大学,博洛尼亚,意大利)
;
School of Electrical and Electronic Engineering(电气与电子工程学院)
William Han, Tony Chen, Chaojing Duan, Xiaoyu Song, Yihang Yao, Yuzhe Yang, Michael A. Rosenberg, Emerson Liu, Ding Zhao
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Allegheny Health Network(阿勒格尼医疗网络)
;
University of California Los Angeles(加州大学洛杉矶分校)
;
University of Colorado(科罗拉多大学)
;
Allergy and Immunology(过敏与免疫学)
专题命中
VLM训练与架构
:VLM(abstract,abstract_cn);vision-language model(abstract);multimodal large language model(abstract);分类 cs.AI
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
AIDEN:面向视障人士的AI助手设计与初步研究
Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer, Bessie Dominguez-Dager, Alvaro Belmonte-Baeza, Mónica Pina-Navarro, Francisco Morillas-Espejo, Felix Escalona, Miguel Cazorla
机构
*
Institute for Computer Research, University of Alicante(计算机研究所,阿利坎特大学)
;
ETH Zurich(苏黎世联邦理工学院)
机构
*
Faculty of Information Technology, University of Jyväskylä(信息科技学院,于韦斯屈耶大学)
;
National University of Singapore(新加坡国立大学)
;
Lab of Brain-Machine Intelligence, Zhejiang University(脑机智能实验室,浙江大学)
;
School of Computer Science and Technology, Dalian University of Technology(计算机科学与技术学院,大连理工大学)