Will Multi-modal Data Improves Few-shot Learning?
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Project Report
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Project Report
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Comments Accepted by International Workshop on Ophthalmic Medical Image Analysis (OMIA) 2020
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Journal ref IEEE Transactions on Geoscience and Remote Sensing, 52(12): 7708 - 7720, 2014
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Comments This paper is accepted to WACV2021
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI
Comments 8 pages, 6 figures, 3 tables. IEEE Robotics and Automation Letters (RA-L)
专题命中 多模态训练与对齐 :multi-modal(title);分类 eess.AS
Comments An updated version of the ICASSP 2020 paper
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV
专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Comments 8 pages, 4 figures, 4 tables
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL
Comments This article was accepted (15 November 2019) and will appear in the proceedings of ICWSM 2020
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Comments arXiv admin note: text overlap with arXiv:1603.01006
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Accepted for publication in WACV 2020
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Accepted at ICCVW (VOT) 2019
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV
Comments 12 pages, part of proceedings for the NAIS 2019 symposium
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Accepted by The Web Conference (WWW) 2019
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18)
专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI
Comments to be published in Conference on Robot Learning (CoRL), 2017
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments ICME 2013
专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV
Comments Submitted to IROS 2016
学习在多模态大语言模型(MLLMs)中预测中间层注意力以进行视觉token剪枝
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对MLLMs视觉token剪枝的固定层非最优及计算成本高的问题,提出MAP方法,实现仅保留5.56%视觉token时维持97.5%性能,获3.09倍端到端加速。
CIGTSurv:结合局部原型关联与全局特征对齐的临床信息引导三模态生存预测
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
AI总结 本研究针对临床信息未充分利用及多模态异质性问题,提出CIGTSurv框架,结合局部原型关联与全局特征对齐机制,在五个TCGA癌症队列上取得生存预测SOTA性能。
Comments Accepted at MICCAI 2026
scMIR:用于单细胞光学显微镜图像表示的视觉语言基础模型
专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 研究针对单细胞光学显微镜图像分析难题,提出scMIR模型,通过自监督图像重建与文本引导跨模态对齐,在多图像文本对上预训练,在多种复杂任务中表现出色,优于现有方法,具备强泛化能力,能推动高通量表型分析工作流程标准化和自动化。
用于文本引导医学分割的骨干网络与语言引导解耦
机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; NingBo No.2 Hospital(宁波第二医院)
专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 研究针对文本引导医学分割中模型组件紧密耦合问题,提出可转移骨干层次适配器框架BTHA,通过稳定特征级接口、分层监督策略和自适应门控语义引导适配器,有效提升分割效果且计算开销适度。
SpaR3D-MoE:来自稀疏视图的自适应3D空间推理与几何归纳专家混合模型
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; YUKUN Intelligent World(宇琨智能世界)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 研究针对多模态大语言模型在2D与3D表征差距问题,提出SpaR3D-MoE框架,通过自适应时空采样和几何归纳专家混合模型,从稀疏RGB输入实现自适应空间推理,在多个实验中取得最优性能。
Comments Accepted to ECCV 2026
对抗文本噪声与冗余:熵感知的密集视觉令牌剪枝
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 提出熵感知密集剪枝(EADP)框架,通过熵过滤文本噪声并利用子模最大化选择令牌,在严格预算下保留细粒度视觉线索,提升VLM精度-效率权衡。
Comments Accepted to ECCV 2026