MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
MoST:通过模态感知混合专家混合语音和文本
Yuxuan Lou, Kai Yang, Yang You
机构
*
School of Computer Science, National University of Singapore(新加坡国立大学计算机科学学院)
;
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.AI、cs.LG
机构
*
CAS Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(中国科学院人工智能安全重点实验室,计算技术研究所,中国科学院,北京,中国)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)
;
Tsinghua University, Beijing, China(清华大学,北京,中国)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.CV
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
用于解释大语言模型 verbalized 自信心的有影响力训练数据检索
Yuxi Xia, Loris Schoenegger, Benjamin Roth
机构
*
Faculty of Computer Science, University of Vienna, Vienna, Austria(维也纳大学计算机科学系)
;
UniVie Doctoral School Computer Science, Vienna, Austria(UniVie计算机科学博士学院)
;
Faculty of Philological and Cultural Studies, University of Vienna, Vienna, Austria(维也纳大学语言与文化研究系)
Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method
推动自适应多阶段视频异常推理:一个基准数据集和方法
Chao Huang, Benfeng Wang, Wei Wang, Jie Wen, Li Shen, Wenqi Ren, Yong Xu, Xiaochun Cao
机构
*
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学计算机科学与技术学院(深圳校区))
;
Shenzhen Key Laboratory of Visual Object Detection and Recognition, Harbin Institute of Technology(哈尔滨工业大学深圳视觉目标检测与识别重点实验室)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Compartmentalised Agentic Reasoning for Clinical NLI
临床自然语言推理中的 compartmentalised agentic 推理
Maël Jullien, Lei Xu, Marco Valentino, André Freitas
机构
*
Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)
;
National Biomarker Centre, CRUK-MI, University of Manchester, UK(曼彻斯特大学国家生物标记中心)
;
Idiap Research Institute, Switzerland(日内瓦IDIAP研究所)
;
School of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系)
;
École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(洛桑联邦理工学院(EPFL))
机构
*
Technical University of Munich, Germany(慕尼黑技术大学)
;
University of Strasbourg, France(斯特拉斯堡大学)
;
University of Glasgow, United Kingdom(格拉斯哥大学)
;
Nanyang Technological University, Singapore(南洋理工大学)
;
University of Massachusetts Boston, USA(马萨诸塞大学波士顿分校)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割
Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
机构
*
Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系)
;
Department of Radiology, University of Florida(佛罗里达大学放射学系)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
Department of Medicine, University of Florida(佛罗里达大学医学系)
;
Department of Radiology, UC San Francisco(旧金山大学放射学系)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
释放大型视觉-语言模型的能力以实现道路基础设施的智能感知
Luxuan Fu, Chong Liu, Bisheng Yang, Zhen Dong
机构
*
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS)(信息工程测绘遥感国家重点实验室)
;
Wuhan University(武汉大学)
;
Hubei Luojia Laboratory(湖北珞珈实验室)
专题命中
视觉定位与Grounding
:vision-language model(title);vision language model(abstract);grounding(abstract);分类 cs.CV
SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery
SVII-3D:利用厘米级3D定位与稀疏街道影像的综合理解,推进道路基础设施库存建设
Chong Liu, Luxuan Fu, Yang Jia, Zhen Dong, Bisheng Yang
机构
*
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS), Wuhan University, Wuhan 430079, China(信息工程测绘遥感国家重点实验室(LIESMARS),武汉大学)
;
Research Institute Ltd, Chengdu 610000, China(四川省公路规划设计研究有限公司)
机构
*
College of Computer Science and Technology, Key Laboratory of Symbolic Computation and Knowledge Engineering, Ministry of Education, Jilin University(吉林大学计算机科学与技术学院,符号计算与知识工程重点实验室,教育部,吉林大学)
;
College of Software, Jilin University(吉林大学软件学院)
;
Facemind Group(FaceMind集团)
;
School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)
机构
*
School of Biomedical Engineering, Southern Medical University(生物医学工程学院,南方医科大学)
;
School of Biomedical Engineering, Shanghai Jiaotong University(生物医学工程学院,上海交通大学)
;
Department of Electronic Engineering, Chinese University of Hong Kong(电子工程系,中国香港大学)
;
Faculty of Dentistry, The University of Hong Kong(牙科学院,香港大学)
;
Department of Nuclear Medicine, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine(核医学科,广州中医药大学第二附属医院)
;
PET Center, Department of Nuclear Medicine, Guangdong Provincial People’s Hospital, Southern Medical University(PET中心,核医学科,广东省人民医院,南方医科大学)
;
Department of Nuclear Medicine, Nanfang Hospital, Southern Medical University(核医学科,南芳医院,南方医科大学)
;
Division of Nuclear Medicine and Molecular Imaging, Geneva University Hospitals(核医学与分子影像学部,日内瓦大学医院)
;
Departments of Radiology, Physics, and Biomedical Engineering, The University of British Columbia(放射学、物理和生物医学工程系,不列颠哥伦比亚大学)
;
Medical Artificial Intelligence Laboratory, Westlake University(医学人工智能实验室,西湖大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV