机构
*
College of Software, Nankai University, China(南开大学软件学院)
;
The University of Hong Kong, China(香港大学)
;
NKIARI, Shenzhen Futian, China(NKIARI,深圳福田,中国)
;
AAIS & VCIP, Nankai University, China(AAIS与VCIP,南开大学,中国)
;
A*STAR Institute of Advanced Intelligence and Computing, Singapore(新加坡A*STAR先进智能与计算研究所)
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
关注、变换或静默:面向高效多模态大语言模型推理的算子级视觉跳跃
Zhaoyang Luo, Runmin Dong, Miao Yang, Fan Wei, Yushan Lai, Bin Luo, Haohuan Fu
机构
*
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
Sun Yat-sen University(中山大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
;
Tsinghua University(清华大学)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Korea University(高丽大学)
;
Upstage AI
;
Kyung Hee University(庆熙大学)
;
KAIST(韩国科学技术院)
;
Hanyang University College of Medicine(汉阳大学医学院)
;
AIGEN Sciences
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
基于LLM增强优化的无人机低空经济网络高效机载视觉-语言推理
Yang Li, Ruichen Zhang, Yinqiu Liu, Guangyuan Liu, Abbas Jamalipour, Xianbin Wang, Dong In Kim
机构
*
College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院、新加坡国立科技大学)
;
The University of Sydney, Sydney, Australia(悉尼大学、澳大利亚悉尼)
;
Department of Electrical and Computer Engineering, Western University, London, Canada(电气与计算机工程系、西方大学、加拿大伦敦)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University, South Korea(电气与计算机工程系、全州大学、韩国)
机构
*
Computer Aided Medical Procedures (CAMP)(计算机辅助医疗程序)
;
TU Munich, Germany(慕尼黑工业大学,德国)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Munich, Germany(慕尼黑,德国)
;
Zhongshan Hospital, Fudan University, China(复旦大学中山医院)
;
The University of Hong Kong, Hongkong, China(香港大学,香港,中国)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
迈向具有空间定位病变证据的临床可解释性眼科VQA
Xingyue Wang, Bo Liu, Meng Wang, Zhixuan Zhang, Chengcheng Zhu, Huazhu Fu, Jiang Liu
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系)
;
The Hong Kong Polytechnic University(香港理工大学)
;
National University of Singapore(新加坡国立大学)
;
University of Washington(华盛顿大学)
;
Institute of High Performance Computing, Agency for Science, Technology and Research(科技研究局高性能计算研究所)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
EVE: 通过可执行视觉变换实现多模态大语言模型的可验证自进化
Yongrui Heng, Chaoya Jiang, Han Yang, Shikun Zhang, Wei Ye
机构
*
National Engineering Research Center for Software Engineering, Peking University(软件工程国家级工程研究中心,北京大学)
;
School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子及计算机工程学系)
;
Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程学系)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所)
;
Center for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
Department of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学系)
;
College of Computing and Data Science (CCDS), Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机科学学院)