Uni-Hema: Unified Model for Digital Hematopathology
机构 * Information Technology University of Punjab(旁遮普信息科技大学) ; Chughtai Lab(楚格塔实验室)
专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Information Technology University of Punjab(旁遮普信息科技大学) ; Chughtai Lab(楚格塔实验室)
专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
机构 * Huazhong University of Science and Technology(华中科技大学) ; Zhongguancun Academy(中关村学院) ; East China Normal University(华东师范大学) ; Zhengzhou University(郑州大学) ; Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
基于CLIP的层次语义树锚定的类增量学习
机构 * School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) ; State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
AI总结 HASTEN通过层次语义树锚定方法,有效解决CLIP类增量学习中的灾难性遗忘问题,提升模型对层次化类别结构的保持能力。
多文本引导的少样本语义分割
机构 * State Key Laboratory of Electromechanical Integrated Manufacturing of High-Performance Electronic Equipments(高性能电子设备机电一体化制造国家重点实验室) ; Center for Complex Systems(复杂系统中心) ; School of Mechano-Electronic Engineering(机械电子工程学院) ; Xidian University(西安电子科技大学)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
AI总结 MTGNet通过融合多文本提示提升少样本语义分割性能,采用多文本先验细化、文本锚点特征融合和前景置信度加权注意力模块,有效增强分割鲁棒性与语义一致性。
机构 * Nanjing University(南京大学) ; Technical University of Munich(慕尼黑技术大学) ; Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍兹中心)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
机构 * Institute of Exact and Natural Sciences(精确与自然科学研究所) ; Federal University of Pará(巴西亚马逊联邦大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Comments 23 pages, 7 figures
机构 * University of Oxford(牛津大学) ; Vienna University of Economics and Business(维也纳经济与商业大学) ; Amazon(亚马逊公司)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments Oral at RecSys Gen AI for E-commerce 2025
机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) ; School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) ; Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) ; Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) ; William and Mary(威廉与玛丽学院) ; Harbin Institute of Technology(哈尔滨工业大学) ; UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)
专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS
Comments Published on IEEE TPAMI
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025
机构 * Department of Software \& Microelectronics Peking University Beijing, China ; Department of Software \& Microelectronics Peking University Beijing, China hangli\ ; Department of Software \& Microelectronics Peking University Beijing, China zehua\ ; Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ ; Economic Law School China University of Political Science ; School of Computer Science Peking University Beijing, China
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
机构 * Beijing Laboratory of Advanced Information Networks, Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications(北京先进信息网络实验室、网络系统架构与收敛重点实验室、信息与通信工程学院、北京邮电大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学)
专题命中 视频多模态 :multimodal(title);分类 cs.AI
机构 * IWR-Bench Team(IWR-Bench团队)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
HV-Attack:多模态检索增强生成的分层视觉攻击
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Sun Yat-Sen University(中山大学) ; Singapore Management University(新加坡管理学院)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出了一种分层视觉攻击方法,通过在图像输入中添加不可察觉扰动,破坏多模态检索增强生成系统的检索和生成性能。
机构 * Department of Computer Science, University of Kaiserslautern-Landau (RPTU)(科斯拉尔特伦大学计算机科学系) ; SDS, German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) ; INRAE, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) ; CIRAD, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) ; INRIA, EVERGREEN, University of Montpellier(蒙彼利埃大学)
专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted at the Machine Learning journal, CfP: Discovery Science 2024
Journal ref Machine Learning 114, 279 (2025)
机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) ; Technical University of Munich(慕尼黑技术大学) ; Visual Geometry Group, University of Oxford(牛津大学视觉几何组)
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024
机构 * HOUMO AI
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments This paper is acceptted by AAAI 2026
专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI
Comments 9 pages, 5 figures
机构 * Queen Mary University of London(伦敦大学玛丽女王学院) ; Institut Jožef Stefan(乔泽夫·斯蒂芬研究所)
专题命中 跨模态检索 :image-text(abstract);分类 cs.CL
Comments Accepted at IJCNLP-AACL 2025
专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract)
动态系统与扩散策略耦合下的理论闭环稳定性边界
机构 * Department of Mechanical Engineering, Universite de Sherbrooke(机械工程系, Sherbrooke 大学) ; Department of Electronic and Computer Engineering, Universite de Sherbrooke(电子与计算机工程系, Sherbrooke 大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.AI
AI总结 本文研究了动态系统与扩散策略耦合下的闭环稳定性边界,提出了一种更快的模仿学习框架和基于演示方差的稳定性判断指标。
Comments 5 pages, 3 figures
机构 * University of Michigan-Ann Arbor(密歇根大学安娜堡分校)
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 多模态生成 :multimodal(abstract)
专题命中 多模态生成 :multimodal(abstract)
Comments 8 pages, 5 figures. Submitted for publication
专题命中 多模态生成 :cross-modal(abstract)
专题命中 多模态生成 :multimodal(abstract)
Comments AAAI2026
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
机构 * Waseda University(早稻田大学) ; NII(日本国立信息机构) ; NII LLMC Tokyo, Japan(日本国立信息机构LLMC东京)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments 17pages, 8 figures
机构 * University of California(加州大学) ; Dana-Farber Cancer Institute(达纳-法伯癌症研究所) ; Massachusetts General Hospital(麻省总医院) ; St. Jude Children’s Research Hospital(圣犹大儿童研究医院) ; Brigham and Women’s Hospital & Harvard Medical School(布里法伦医院及哈佛医学院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Proceedings of the 5th Machine Learning for Health (ML4H) Symposium
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments 11 pages, 5 figures
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments 8 pages, 1 figure