Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI
机构 * Lingran Song, Yucheng Zhou, Jianbing Shen(作者)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments AAAI 2026
机构 * Department of EECS University of Arkansas(电子工程与科学系 亚拉荷加大学) ; Department of CS Baylor University(计算机科学系 基尔默大学)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV
Comments This version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission
机构 * Nova School of Business and Economics(诺瓦商业与经济学院) ; Nova School of Business(诺瓦商业学院)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accetped at IEEE BigData 2025, 10 pages, 5 figures, 3 tables
专题命中 图文多模态 :multimodal(title,abstract)
机构 * Technical University of Munich (TUM)(慕尼黑技术大学) ; TUM University Hospital(慕尼黑技术大学医院) ; German Heart Center TUM University Hospital(慕尼黑技术大学医院德国心脏中心) ; Department of Radiology(放射科) ; Klinikum rechts der Isar TUM University Hospital(慕尼黑技术大学医院右岸诊所) ; Uniklinik RWTH Aachen(亚琛工业大学医院) ; HOPPR IL USA(HOPPR美国) ; University of Oxford(牛津大学) ; Imperial College London(伦敦帝国学院)
专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL
Comments Accepted to ML4H 2025 Proceedings
机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) ; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) ; Pengcheng Laboratory(鹏城实验室) ; University of Chinese Academy of Sciences(中国科学院大学) ; Chinese Medicine Guangdong Laboratory(广东中医药实验室) ; Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha 410073, China(大数据与决策实验室,国防科技大学,长沙410073,中国)
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI
Comments 15 page, 9 figures, published to PRCV
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments Accepted by AAAI 2026 as an Oral Presentation (13 pages, 7 figures, 7 tables)
Journal ref AAAI2026
机构 * esieaLab, ETIS Laboratory(esiea实验室,ETIS实验室) ; ESIEA, University of CY Cergy(ESIEA,CY塞克大学) ; ETIS Laboratory, CNR1S, UMR8051(ETIS实验室,CNR1S,UMR8051)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 5 pages, 3 figures, ICTAI 2025
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM
Comments 10 pages
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL、cs.AI
Comments Accepted at the Ninth Conference on Machine Translation (WMT24), co-located with EMNLP 2024
Journal ref https://aclanthology.org/2024.wmt-1.81/
机构 * Heinrich Heine University of Düsseldorf(海因里希-海涅大学杜塞尔多夫分校)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments Accepted in WACV 2026. Code in https://github.com/HHU-MMBS/clustermine_wacv_official 9 Tables, 11 Figures
机构 * Department of Cyber-Physical Systems, Clark Atlanta University(克雷克阿特拉大学计算机物理系统系) ; Siemens Corporation(西门子公司)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、cs.MM、eess.AS
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI
机构 * PwC US(普华永道美国)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL
Comments 15 pages, 4 figures
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.MM
机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments 14 pages, 6 figures
机构 * Massachusetts Institute of Technology(麻省理工学院)
专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
机构 * The Hong Kong Polytechnic University(香港理工大学) ; ARC Lab, Tencent PCG(腾讯PCG ARC实验室) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; vivo Mobile Communication Co.(vivo移动通信公司) ; MindWingman Technology (Shenzhen) Co., Ltd.(深圳MindWingman技术有限公司)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments NeurIPS 2025 Camera Ready. Project Page: https://polyu-chenlab.github.io/unipixel/
机构 * Tsinghua University(清华大学) ; Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) ; ZTE Corporation(中兴通讯)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI
机构 * University of California, Merced(加州大学默塞德分校)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments Preliminary version, 19 pages
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments AAAI 2026 (Oral presentation)
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学) ; Shanghai Jiao Tong University(上海交通大学) ; Wuhan University(武汉大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(abstract)
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI
专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV