xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
机构 * Salesforce AI Research(Salesforce AI研究院) ; Intel Labs(英特尔实验室) ; University of Washington(华盛顿大学)
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Salesforce AI Research(Salesforce AI研究院) ; Intel Labs(英特尔实验室) ; University of Washington(华盛顿大学)
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
机构 * POSTECH ; University of California, Berkeley(加州大学伯克利分校)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2025 Main Conference
专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM
机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)
专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV
Comments Technical Report. Project Page: https://klingavatar.github.io/
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
机构 * Virginia Tech(维吉尼亚理工大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI
机构 * ADAPT Centre(ADAPT中心) ; Dublin City University (DCU)(都柏林城市大学) ; Trinity College Dublin (TCD)(三一学院都柏林)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments Accepted at WMT2025 (ENNLP) for oral presented
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
专题命中 音频语音多模态 :multimodal(abstract)
Comments 11 pages, 6 figures, to appear in IEEE TVCG (VIS 2024); correct figure
Journal ref IEEE Transactions on Visualization and Computer Graphics, 31(1), 2025, 941-951
机构 * Multimodal Intelligence Lab, Department of Computer Science University of Exeter Exeter, UK(埃克塞特大学计算机科学系多模态智能实验室) ; School of Computer Science University of Birmingham Birmingham, UK(伯明翰大学计算机科学学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Laboratoire Informatique d'Avignon, Avignon University, France(阿维尼翁信息实验室,阿维尼翁大学,法国)
专题命中 视频多模态 :multimodal(title,abstract)
Comments Paper accepted at ICPRAM 2025
机构 * Graduate School of AI, KAIST(韩国国立庆熙大学人工智能研究生院) ; Graduate School of CT, KAIST(韩国国立庆熙大学CT研究生院)
专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS
Comments Accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
机构 * Nanjing University of Science and Technology(南京理工大学) ; The Hong Kong University of Science and Technology(香港科学大学) ; Nanjing Forestry University(南京林业大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
机构 * School of Artificial Intelligence and Robotics(人工智能与机器人学院) ; National Engineering Research Center for Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) ; Hunan University(湖南大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted to IEEE/ASME Transactions on Mechatronics
Journal ref IEEE/ASME Transactions on Mechatronics, 2025
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV
Comments ICCV 2025 MARS2 Workshop and Challenge "Multimodal Reasoning and Slow Thinking in the Large Model Era: Towards System 2 and Beyond''
机构 * Dali University(大理大学) ; Bandırma Onyedi Eylül University(巴尔迪马十一点大学)
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV
机构 * University of Naples Federico II(那不勒斯费迪里奇二世大学) ; Northwestern University(西北大学)
专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI
Journal ref SIGIR '25: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025
专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM
Journal ref Liang B, Wang B, Bai Z, et al. Reply with Sticker: New Dataset and Model for Sticker Retrieval[J]. IEEE Transactions on Audio, Speech and Language Processing, 2025
专题命中 跨模态检索 :multimodal(abstract)
Comments Forthcoming in Frontiers in Education (FIE 2025), Nashville, Tennessee, USA, Nov 2-5, 2025
专题命中 跨模态检索 :multimodal(abstract)
机构 * Zhejiang University(浙江大学)
专题命中 多模态生成 :multimodal(title);MLLM(abstract);image-text(abstract);分类 cs.CV
机构 * Shenzhen Technology University(深圳科技大学) ; University of Washington(华盛顿大学) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 13 pages,12 figures
机构 * Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)
专题命中 多模态生成 :multi-modal(title);分类 cs.AI
Comments Major changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section
机构 * Future Robotics Organization, Waseda University(早稻田大学未来机器人组织) ; Microsoft Research Asia(微软亚洲研究院) ; Department of Modern Mechanical Engineering, Waseda University(早稻田大学现代机械工程系) ; Artificial Intelligence Laboratories, Fujitsu Limited(Fujitsu 人工智能实验室) ; Center for Technology Innovation - Controls and Robotics, Research & Development Group, Hitachi, Ltd.(富士通技术研发集团技术创新中心 - 控制与机器人) ; Faculty of Science and Engineering, Waseda University(早稻田大学工学部) ; National Institute of Advanced Science and Technology(国家先进科学研究院)
专题命中 多模态生成 :multimodal(title)
Comments Humanoids2025
机构 * Korea University(韩国大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV
机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) ; Technical University Berlin(柏林技术大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CL
Comments Presented and published at BioCreative IX
机构 * Shenzhen University(深圳大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; Zhejiang University(浙江大学)
专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI
Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465
专题命中 多模态评测 :multi-modal(title,abstract);MLLM(abstract);分类 cs.MM
机构 * Singapore University of Technology and Design(新加坡科技设计大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments 27 pages, 8 figures, EMNLP 2025 Findings
机构 * Robotics Research Centre of the School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore(南洋理工大学机械与航空航天工程学院机器人研究中心) ; College of Information and Control Engineering, Xi’an University of Architecture and Technology, Xi’an, China(西安建筑科技大学信息与控制工程学院)
专题命中 多模态评测 :cross-modal(title,abstract)
Comments 9 pages, 6 figures