Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
机构 * Shanghai Jiao Tong University(上海交通大学) ; UC Berkeley(伯克利大学) ; UNC, Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Shanghai Jiao Tong University(上海交通大学) ; UC Berkeley(伯克利大学) ; UNC, Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)
机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) ; University of Surrey(Surrey大学) ; School of Computer Science, Wuhan University(武汉大学计算机学院) ; School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) ; College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) ; Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables
机构 * Sangho Lee 1,2(Sangho Lee 教授) ; Il Yong Chun 1,3(Il Yong Chun 教授) ; Hogun Park 1(Hogun Park 教授)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted to the AAAI 2025 Main Technical Track. This is an extended version of the original submission
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Purdue University(普渡大学) ; University at Albany(阿尔巴尼大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted to AAAI 2025
机构 * Faculty of Engineering and Architecture, IDLab-AIRO, Ghent University – imec, Technologiepark 126, 9052 Gent, Belgium(工程与建筑学院,IDLab-AIRO,根特大学–imec,Technologiepark 126,9052 Gent,比利时)
专题命中 音频语音多模态 :multimodal(title,abstract)
专题命中 音频语音多模态 :multimodal(title,abstract)
Comments 18 pages, 5 figures, 4 tables
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Northeastern University(东北大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Chinese Academy of Sciences(中国科学院大学) ; Tsinghua University(清华大学) ; Sun Yat-sen University(中山大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments 26 pages, 23 figures, the code is available at \url{https://github.com/DabDans/AudioMarathon}
机构 * Stanford University(斯坦福大学) ; National University of Singapore(新加坡国立大学) ; Sound Speech and Hearing Clinic(语音与听力诊所)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments EMNLP 2025 Oral Presentation
机构 * McGill University(麦吉尔大学) ; CRBLM ; Mila Quebec AI Institute(魁北克人工智能研究所) ; Université de Montréal(蒙特利尔大学) ; Nouvelle Voix(新声音) ; Montreal Neurological Institute(蒙特利尔神经科学研究所) ; Concordia University(Concordia大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to SMASH 2025
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted to AAAI 2025
机构 * Southern University of Science and Technology(南方科技大学) ; Tencent Inc.(腾讯公司)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by ACCV 2022
机构 * ByteDance(字节跳动)
专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to AAAI 2025
机构 * Université Paris Cité(巴黎Cité大学) ; University of Oxford(牛津大学) ; University of Turin(都灵大学) ; Universität Leipzig(莱比锡大学) ; Max Planck School of Cognition(马克斯·普朗克认知科学学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Shanghai Jiao Tong University(上海交通大学) ; Alibaba group(阿里巴巴集团)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(abstract)
Comments RecSys 2025 Industry Track
机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) ; King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
Comments EMNLP Mains 2025
机构 * Stony Brook University(史坦尼·布鲁克大学) ; University Paris 1 Pantheon-Sorbonne(巴黎第一大学)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI
专题命中 跨模态检索 :multi-modal(title);分类 cs.AI
Comments Accepted in NeurIPS 2025
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Xiaohongshu Inc.(小红书公司)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * CUHK(香港中文大学) ; HKUST(香港科技大学) ; HKU(香港大学) ; ByteDance Inc(字节跳动公司)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Shanghai Innovation Institute(上海创新研究院) ; Nanjing University(南京大学) ; The University of Sydney(悉尼大学) ; Shanghai Jiao Tong University(上海交通大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
Comments 33 pages, 13 figures, 10 tables
专题命中 多模态生成 :cross-modal(abstract)
Comments 13 main pages, 5 figures, 2 tables
机构 * Baidu Inc.(百度公司)
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Accepted by AAAI2025
机构 * Department of Computer Science University of Manchester(曼彻斯特大学计算机科学系) ; Department of Advanced Information Technology Kyushu University(九州大学先进信息技术系)
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
Comments Accepted as an oral presentation at the EMNLP 2025 Workshop on Machine Translation (WMT)
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments Work in progress
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
Comments 7 Pages, 3 Figures, The Dream2Image dataset is openly available on Hugging Face at: https://huggingface.co/datasets/opsecsystems/Dream2Image
机构 * Singapore Management University(新加坡管理大学) ; Fudan University(复旦大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments Project website: https://yxg1005.github.io/GaslightingNegationAttacks/
机构 * Worcester Polytechnic Institute(沃斯特理工学院)
专题命中 多模态评测 :multimodal(title,abstract)
机构 * Italian Institute of Germanic Studies (IISG)(意大利德语研究学院) ; University of Bologna(博洛尼亚大学)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments The First Workshop on Natural Language Processing and Language Models for Digital Humanities (LM4DH 2025). RANLP 2025