Cross-Modal Attention Guided Unlearning in Vision-Language Models
专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV
机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) ; Hong Kong Polytechnic University(香港理工大学) ; Zhuhai Campus of Sun Yat-sen University(中山大学珠海校区) ; University of Adelaide(阿德莱德大学) ; Peking University(北京大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project Page: https://zju-real.github.io/SpatialLadder/ Code: https://github.com/ZJU-REAL/SpatialLadder
机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.AI
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) ; TU Munich(慕尼黑技术大学)
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 63 pages, 29 tables, and 47 figures
机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS
Comments 6 pages,2 figures,accepted by ACM Multimedia Asia 2025
机构 * Shanghai Jiao Tong University(上海交通大学) ; Ant Group(蚂蚁集团)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments submitted to ICASSP 2026
机构 * Bar Ilan University(巴伊兰大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
机构 * University of Central Florida(中央佛罗里达大学) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) ; Stanford University(斯坦福大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
机构 * Tsinghua University(清华大学) ; Huazhong University of Science and Technology(华中科技大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
机构 * Nanjing University(南京大学) ; ByteDance(字节跳动) ; Nankai University(南开大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Shanghai Jiao Tong University(上海交通大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Bilibili Inc.(哔哩哔哩公司)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * State Key Laboratory of Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学) ; Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) ; Tsinghua University(清华大学) ; ModelTC
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)
专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV
机构 * Intelligent Creation Team, ByteDance(字节跳动智能创作团队)
专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV
机构 * City University of Hong Kong(香港城市大学) ; WeChat, Tencent Inc(微信、腾讯公司) ; Manycore Tech Inc(很多核科技公司)
专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV
Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video
机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project Page: https://showlab.github.io/Paper2Video/
机构 * Shanghai Jiaotong University(上海交通大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; East China Normal University(华东师范大学) ; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV
Comments 25pages,20figures
机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
Comments Website with code: https://guowei-zou.github.io/dm1/
机构 * Microsoft(微软公司) ; Massachusetts General Hospital, Harvard University(哈佛大学麻省总医院) ; Emory University(埃默里大学)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
专题命中 多模态生成 :multi-modal(abstract)
机构 * School of Computing and Communications(计算与通讯学院) ; Lancaster Medical School(兰卡斯特医学学院)
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract,comments);分类 cs.CV、cs.AI
Comments ICCV 2025 (PHAROS-AFE-AIMI: Adaptation, Fairness, and Explainability in Medical Imaging). 8 pages, 5 figures, 4 tables. Keywords: multi-modal, multimodal, prototype learning, explainable AI, interpretable models, case-based reasoning, medical imaging, DEXA, bone health, osteoporosis, osteopenia, diagnosis, classification, clustering
机构 * West China Biomedical Big Data Center, West China Hospital, SCU(西昌生物医学大数据中心、西昌医院、SCU) ; NTU ; USTC ; TeleAI, China Telecom(TeleAI、中国电信) ; BUPT ; Tsinghua University(清华大学) ; HKUST(Guangzhou)(HKUST(广州))
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
机构 * Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究生院) ; RIKEN Center for Advanced Intelligence Project (AIP), RIKEN(日本理化学研究院先进智能项目中心) ; Department of Photogrammetry and Remote Sensing, ETH Zürich(苏黎世联邦理工学院测绘与遥感系) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Microsoft Research(微软研究院) ; Cornell University(康奈尔大学) ; University of Cambridge(剑桥大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Ericsson Research(爱立信研究)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Oral presentation at ICIP 2025