Generating Accurate and Detailed Captions for High-Resolution Images
Hankyeol Lee, Gawon Seo, Kyounggyu Lee, Dogun Kim, Kyungwoo Song, Jiyoung Jung
机构
*
Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系)
;
Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)
;
Department of Applied Statistics, Yonsei University(延世大学应用统计系)
专题命中
图文多模态
:multimodal(abstract);分类 cs.CV、cs.AI
CommentsWork conducted in 2024; released for archival purposes
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
Jiarong Du, Zhan Jin, Peijun Yang, Juan Liu, Zhuo Li, Xin Liu, Ming Li
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems, Digital Innovation Research Center, Duke Kunshan University(多模态智能系统苏州市级重点实验室、杜克昆山大学数字创新研究中心)
;
Hardware Engineering System, OPPO(OPPO硬件工程系统)
UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
Yudong Yang, Xiaokang Liu, Shaofeng zhao, Rongfeng Su, Nan Yan, Lan Wang
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, China(深圳先进技术研究院,中国科学院,中国)
;
Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences, China(生物医学成像科学与系统重点实验室,中国科学院,中国)
Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation
Elena Mulero Ayllón, Linlin Shen, Pierangelo Veltri, Fabrizia Gelardi, Arturo Chiti, Paolo Soda, Matteo Tortora
机构
*
Unit of Artificial Intelligence and Computer Systems, Università Campus Bio-Medico di Roma, Italy(人工智能与计算机系统单位,罗马大学生物医学校园,意大利)
;
College of Computer Science and Software Engineering, Shenzhen University, China(计算机科学与软件工程学院,深圳大学,中国)
;
Dept. of Computer Engineering, Modeling, Electronic and System Engineering, University of Calabria, Italy(计算机工程、建模、电子与系统工程系,卡拉布里亚大学,意大利)
;
IRCCS San Raffaele Hospital, Italy(圣拉斐拉医院,意大利)
;
Faculty of Medicine, Vita-Salute San Raffaele University, Italy(医学学院,维塔-桑拉斐拉大学,意大利)
;
Dept. of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering, Umeå University, Sweden(诊断与干预系,辐射物理,生物医学工程,乌梅大学,瑞典)
;
Dept. of Naval, Electrical, Electronics and Telecommunications Engineering, University of Genoa, Italy(海军、电子、电子与电信工程系,热那亚大学,意大利)
机构
*
Hong Kong Baptist University(香港 Baptist 大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
National University of Singapore(新加坡国立大学)
Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen
机构
*
IRIT, University of Toulouse, France(IRIT,图卢兹大学,法国)
;
Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)
;
CNRS, IRIT, France(CNRS,IRIT,法国)
LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
Magdalena Lederbauer, Siddharth Betala, Xiyao Li, Ayush Jain, Amine Sehaba, Georgia Channing, Grégoire Germain, Anamaria Leonescu, Faris Flaifil, Alfonso Amayuelas, Alexandre Nozadze, Stefan P. Schmid, Mohd Zaki, Sudheesh Kumar Ethirajan, Elton Pan, Mathilde Franckel, Alexandre Duval, N. M. Anoop Krishnan, Samuel P. Gleason
机构
*
Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)
;
Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢赫大学)