机构
*
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
;
College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院)
;
School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室)
;
School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室)
;
Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
Li Zhou, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li, Haizhou Li
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Research Institute of Big Data(深圳大数据研究院)
;
Chengdu Technological University(成都理工大学)
;
University of Copenhagen(哥本哈根大学)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
Pei Liu, Hongliang Lu, Haichao Liu, Haipeng Liu, Xin Liu, Ruoyu Yao, Shengbo Eben Li, Jun Ma
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
Li Auto Inc.
;
the School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung
机构
*
Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院)
;
The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
;
National University of Singapore(新加坡国立大学)
;
CVTE
;
Sun Yat-sen University(孙中山大学)
;
Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科)
;
Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科)
;
School of Computer Science, Peking University(北京大学计算机科学系)
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
Hongyu Ge, Longkun Hao, Zihui Xu, Zhenxin Lin, Bin Li, Shoujun Zhou, Hongjin Zhao, Yihang Liu
机构
*
The Hong Kong University of Sciences and Technology, Guangzhou(香港科学与技术大学(广州))
;
Beihang University(北航大学)
;
Shandong University(山东大学)
;
Hubei University(湖北大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
;
Australian National University(澳大利亚国立大学)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
Saehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi, Sungroh Yoon
机构
*
Department of Electrical and Computer Engineering, Seoul National University(电气与计算机工程系,首尔国立大学)
;
Interdisciplinary Program in Artificial Intelligence, Seoul National University(人工智能交叉学科项目,首尔国立大学)
;
Adobe Research(Adobe研究)
专题命中
视觉问答
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV