D2 Pruning: Message Passing for Balancing Diversity and Difficulty in Data Pruning
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 17 pages (Our code is available at https://github.com/adymaharana/d2pruning)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 17 pages (Our code is available at https://github.com/adymaharana/d2pruning)
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments arXiv admin note: text overlap with arXiv:2210.09338 by other authors
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments CVPR 2023 (Highlighted Paper). Website: https://imagebind.metademolab.com/ Code/Models: https://github.com/facebookresearch/ImageBind
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)
Comments 5 pages, accepted to The Industry Track of the Web Conference 2023
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published at ICLR 2023
专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 5 pages, 2 figures, accepted by 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022)
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted as Student Abstract at AAAI-22
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted for publication at the 1st Multilingual Representation Learning workshop (MRL 2021) co-located with EMNLP 2021. 15 pages, 8 figures, 6 tables
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments 17 pages, 3 figures
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)
Comments 13 Pages, 6 Figures and 10 Tables
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments Accepted to ACM Transactions on Multimedia Computing Communications and Applications (ACM TOMM)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 3 pages including references, Accepted at the ICCV 2019 Workshop - 'Linguistics Meets Image and Video Retrieval' (received Best Paper Award)
专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments *Ayush Jaiswal and Ekraam Sabir contributed equally to the work in this paper
Journal ref In Proceedings of the 2017 ACM on Multimedia Conference, pp. 1465-1471. ACM, 2017
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to BMVC'16
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)
MAGMaR 2026 共享任务结果
机构 * Johns Hopkins University(约翰霍普金斯大学) ; OpenAI ; University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) ; Air Force Research Laboratory(空军研究实验室) ; Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心) ; University of Amsterdam(阿姆斯特丹大学) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.CL
AI总结 本文介绍MAGMaR 2026共享任务的结果,包括视频检索和基于检索视频的生成任务,所有提交系统均超越去年基线。
Comments Findings of the 2nd workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR); Resources at this url: https://github.com/rekriz11/MAGMAR_2026
草图与文本协同:融合结构轮廓和描述属性用于细粒度图像检索
机构 * Xidian University First Aircraft Design Institute(西安电子科技大学第一飞机设计院)
专题命中 跨模态检索 :cross-modal(abstract,comments);分类 cs.CV、cs.AI
AI总结 本文提出STBIR框架,通过融合草图的结构轮廓与文本的色彩纹理信息,提升细粒度图像检索性能,采用课程学习、特征空间优化和多阶段跨模态对齐机制,实验验证其优于现有方法。
Comments Image Retrieval, Hand-drawn Sketch, Multi-stage Cross-modal Feature Alignment
机构 * ByteDance(字节跳动) ; S-Lab, NTU(NTU的S实验室)
专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.CL
Comments Code: https://github.com/EvolvingLMMs-Lab/multimodal-search-r1
专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.AI
Comments SIGIR 2024 Workshop on Multimodal Representation and Retrieval (MRR 2024)
基于结构化上下文推理提升基于知识的视觉问答性能
机构 * School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) ; School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) ; Graduate School of Advanced Science and Engineering, Hiroshima University(广岛大学先进科学与工程研究生院)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM
AI总结 本文提出SCoRe框架,通过上下文获取、选择、压缩三阶段处理多模态知识,在OK-VQA和A-OKVQA基准上性能优于现有最优方法,提升了基于知识的视觉问答效果。
Comments Accepted by ICME 2026. Source code is available at https://github.com/WISLab-GDUT/SCoRe
双流跨锚校正:接地长文本描述与对象级锚点的域限制
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL
AI总结 针对多模态大语言模型的长文本描述对象幻觉问题,本文提出双流跨锚校正方法,通过耦合感知流与认知流提升精度,在长文本场景下实现最优性能,且存在域条件性限制。
开放词汇全景分割与检索增强
专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL
AI总结 本文提出RetCLIP方法,通过检索增强提升开放词汇全景分割的性能,实现对未见类别的有效分割。
Journal ref IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026)
无校正控制的基础:大语言模型的真值追踪剖面
机构 * Humber Polytechnic(亨伯理工学院) ; University of Toronto(多伦多大学)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI
AI总结 本文研究大语言模型中无校正控制的基础问题,提出路径剖面概念以分析真值追踪,指出纯文本模型继承的模式可提供衍生可应答性,不同方法对任务的真值追踪改进可能与表面改进不一致。
Comments 24 pages, 1 figure, 1 table. A six-page methodological supplement, reproducible R script, and constructed data are included as ancillary files
CMCNet:将超声图像嵌入与文本TI-RADS表示对齐以用于细粒度甲状腺分类
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本研究构建含600个甲状腺结节的STN数据集,提出CMCNet模型,通过中心间隔对比损失对齐图像与文本嵌入,实现细粒度甲状腺分类,性能优于多类基线方法。