Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
机构 * Norwegian University of Science and Technology (NTNU)(挪威科学技术大学) ; MOBAI AS(MOBAI公司)
专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Norwegian University of Science and Technology (NTNU)(挪威科学技术大学) ; MOBAI AS(MOBAI公司)
专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL
机构 * South China Normal University(华南师范大学) ; Shenzhen Polytechnic University(深圳职业技术大学)
专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV
Comments Accepted by ICCV2025
机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments 11 pages, 1 figure
机构 * School of Computer Science and Artificial Intelligence, Wuhan Textile University(武汉纺织大学计算机科学与人工智能学院) ; School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * University of Freiburg(弗赖堡大学) ; deepset
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
Comments accepted at GCPR 2025
专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Under Review
机构 * Center For Artificial Intelligence and Data Science, University of Würzburg(人工智能与数据科学中心,乌尔姆大学) ; Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) ; Mila, McGill University and Canada CIFAR AI Chair(Mila,麦吉尔大学及加拿大CIFAR人工智能主席)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI
机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; School of Computer Technology and Application, Qinghai University(计算机技术与应用学院,青海大学) ; ByteDance(字节跳动) ; Nanjing University(南京大学) ; Southern University of Science and Technology(南方科技大学) ; Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 34 pages, 10 figures
机构 * Speech and Hearing (SPandH), School of Computer Science, The University of Sheffield, UK(语音与听力(SPandH)、计算机科学学院、谢菲尔德大学、英国) ; South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗兰肯应用科学大学、伊塞尔洛恩、德国)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM
Comments Accepted to SPECOM 2025, 13 pages, 4 figures. To appear in the Proceedings of the 27th International Conference on Speech and Computer (SPECOM) 2025, October 13-14, 2025, Szeged, Hungary
专题命中 音频语音多模态 :multimodal(abstract)
Comments 6 figures, 4 tables
机构 * Applied Artificial Intelligence Institute, Deakin University, Australia(应用人工智能研究所,德金大学,澳大利亚)
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
机构 * Amazon(亚马逊) ; School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学) ; Ben-Gurion University of the Negev(内盖夫本·古里安大学) ; University of Cambridge(剑桥大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
机构 * Medical AI Research Center ( MedARC )(医学人工智能研究中心(MedARC)) ; Baylor College of Medicine(贝勒医学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments Perspective piece on Algonauts 2025 Challenge conclusion
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Northwestern Polytechnical University(西北工业大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
机构 * University of Central Florida(中央佛罗里达大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted to VISION'25 - ICCV 2025 workshop
机构 * University of Central Florida(中央佛罗里达大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Shanghai AI LAB(上海人工智能实验室)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments work in process
机构 * Shanghai Jiao Tong University(上海交通大学) ; Singapore Institute of Technology(新加坡科技学院)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.MM
机构 * 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited ; 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles ; 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute ; 9Noah's Ark Lab 10CG Matrix Technology Limited
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI
Comments Accepted at the 34th ACM International Conference on Information and Knowledge Management (CIKM2025)
机构 * Zhejiang University, China(浙江大学)
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.AI
Comments 8 pages, 5 figures
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Peking University(北京大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Alibaba Group(阿里巴巴集团) ; National University of Defense Technology(国防科技大学) ; Meta
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments 10 pages, 12 figures
机构 * Basic Algorithm Center, PCG, Tencent(腾讯基础算法中心、PCG、腾讯)
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
机构 * Paris-Saclay University(巴黎-萨克雷大学) ; ONERA - The French Aerospace Lab(法国航空航天实验室)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV
Comments Project Page: https://haroldchen19.github.io/PhysHPO-Page/
专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV
机构 * Linqing Chen, Hanmeng Zhong, Wentao Wu, Weilei Wang(作者)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL