arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2504.17990 2025-04-28 cs.CV 57%

From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval

Yabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin, Sanping Zhou, Ming Yang, Le Wang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究所,西安交通大学) Harbin Institute of Technology(哈尔滨工业大学) Ant Group(蚂蚁集团)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04918 2025-04-25 cs.CV 57%

Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity

Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang

机构 * National Sun Yat-sen University(国立中山大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 6 figures, International Conference on Technologies and Applications of Artificial Intelligence (TAAI) Camera Ready

Journal ref Technologies and Applications of Artificial Intelligence, pp. 77-90, Springer, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08548 2025-04-16 cs.GR cs.CV 57%

COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails

Miguel Espinosa, Valerio Marsocci, Yuru Jia, Elliot J. Crowley, Mikolaj Czerkawski

机构 * University of Edinburgh(爱丁堡大学) European Space Agency (ESA)(欧洲空间局) KU Leuven(鲁汶大学) Asterisk Labs(星号实验室)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2025 Workshop MORSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10084 2025-04-15 cs.CV 57%

UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval

Yating Liu, Yaowei Li, Xiangyuan Lan, Wenming Yang, Zimo Liu, Qingmin Liao

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Peng Cheng Laboratory(鹏城实验室) School of ECE, Peking University(北京大学电子与计算机工程学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 7 figures, first submited to IEEE TCSVT on 2024 May. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09302 2025-04-15 cs.AI 57%

Application of Contrastive Learning on ECG Data: Evaluating Performance in Japanese and Classification with Around 100 Labels

Junichiro Takahashi, JingChuan Guan, Masataka Sato, Kaito Baba, Kazuto Haruguchi, Daichi Nagashima, Satoshi Kodera, Norihiko Takeda

机构 * The University of Tokyo Hospital(东京大学医院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 13 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03128 2025-04-07 cs.CV 57%

FontGuard: A Robust Font Watermarking Approach Leveraging Deep Font Knowledge

Kahim Wong, Jicheng Zhou, Kemou Li, Yain-Whar Si, Xiaowei Wu, Jiantao Zhou

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02397 2025-04-04 cs.CV 57%

Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval

Boseung Jeong, Jicheol Park, Sungyeon Kim, Suha Kwak

机构 * POSTECH(浦项科技大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02310 2025-04-04 cs.CL 57%

Improving Harmful Text Detection with Joint Retrieval and External Knowledge

Zidong Yu, Shuo Wang, Nan Jiang, Weiqiang Huang, Xu Han, Junliang Du

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11005 2025-04-03 cs.CV 57%

Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection

Chuhan Zhang, Chaoyang Zhu, Pingcheng Dong, Long Chen, Dong Zhang

机构 * The Hong Kong University of Science and Technology(香港科技大学) AI Chip Center for Emerging Smart Systems (ACCESS)(新兴智能系统AI芯片中心)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments 10 pages, 5 figures, Published as a conference paper at ICLR 2025

Journal ref Proceedings of the 13th International Conference on Learning Representations (ICLR 2025), Paper ID: 4226

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01009 2025-04-02 cs.CV 57%

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

Saarthak Kapse, Pushpak Pati, Srikar Yellapragada, Srijan Das, Rajarsi R. Gupta, Joel Saltz, Dimitris Samaras, Prateek Prasanna

机构 * Stony Brook University(石溪大学) UNC Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14746 2025-04-02 cs.CV 57%

CoVR-2: Automatic Data Construction for Composed Video Retrieval

Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

机构 * École des Ponts(巴黎路桥学院) Univ Gustave Eiffel(Gustave Eiffel大学) CNRS(法国国家科学研究中心) Google DeepMind(谷歌DeepMind) ENS(巴黎高等师范学院) Inria(法国国家信息与自动化研究院)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Appears in TPAMI 2024 (DOI: 10.1109/TPAMI.2024.3463799). Journal extension of the AAAI 2024 conference paper arXiv:2308.14746v3. Project page: https://imagine.enpc.fr/~ventural/covr/

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17007 2025-04-02 cs.CV 57%

Where am I? Cross-View Geo-localization with Natural Language Descriptions

Junyan Ye, Honglin Lin, Leyan Ou, Dairong Chen, Zihao Wang, Qi Zhu, Conghui He, Weijia Li

机构 * Sun Yat-Sen University(中山大学) Shanghai AI Laboratory(上海人工智能实验室) Sensetime Research(商汤科技研究院) Wuhan University(武汉大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07622 2025-04-01 cs.CV 57%

Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval

Junyang Chen, Hanjiang Lai

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Sun Yat-Sen University(中山大学)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments accepted by ICME 2025, this is the full version of paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22064 2025-03-31 cs.AI cs.SY eess.SY 57%

Multi-Task Semantic Communications via Large Models

Wanli Ni, Zhijin Qin, Haofeng Sun, Xiaoming Tao, Zhu Han

机构 * Tsinghua University(清华大学) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心) Beijing University of Posts and Telecommunications(北京邮电大学) University of Houston(休斯顿大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19910 2025-03-26 cs.CV cs.IR 57%

CoLLM: A Large Language Model for Composed Image Retrieval

Chuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, Abhinav Shrivastava

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025. Project page: https://collm-cvpr25.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18016 2025-03-25 cs.CV 57%

Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

Xu Zheng, Ziqiao Weng, Yuanhuiyi Lyu, Lutao Jiang, Haiwei Xue, Bin Ren, Danda Paudel, Nicu Sebe, Luc Van Gool, Xuming Hu

机构 * HKUST(GZ)(香港科技大学(广州)) INSAIT(保加利亚人工智能与数据科学研究所) Sofia University “St. Kliment Ohridski”(索非亚大学) HKUST(香港科技大学) Sichuan University(四川大学) Tinghua University(清华大学) ETH Zurich(苏黎世联邦理工学院) University of Pisa(比萨大学) University of Trento(特伦托大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09313 2025-03-18 cs.CL cs.IR 57%

xVLM2Vec: Adapting LVLM-based embedding models to multilinguality using Self-Knowledge Distillation

Elio Musacchio, Lucia Siciliani, Pierpaolo Basile, Giovanni Semeraro

机构 * University of Bari Aldo Moro(巴里大学阿尔多莫罗分校)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments fix typo in number of tasks in MMEB; fix url for source code; added missing reference to XTD10

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15839 2025-03-18 cs.CV 57%

VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding

Jiaqi Wang, Yifei Gao, Jitao Sang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14594 2025-03-18 cs.CV cs.RO 57%

VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition

Yun-Jin Li, Mariia Gladkova, Yan Xia, Rui Wang, Daniel Cremers

机构 * TU Munich(慕尼黑工业大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Microsoft(微软公司)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Project page https://yunjinli.github.io/projects-vxp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08323 2025-03-12 cs.CL 57%

Towards Scalable and Cross-Lingual Specialist Language Models for Oncology

Morteza Rohanian, Tarun Mehra, Nicola Miglino, Farhad Nooralahzadeh, Michael Krauthammer, Andreas Wicki

机构 * University of Zurich(苏黎世大学) University Hospital Zurich(苏黎世大学医院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02162 2025-03-12 cs.CV cs.LG 57%

X2CT-CLIP: Enable Multi-Abnormality Detection in Computed Tomography from Chest Radiography via Tri-Modal Contrastive Learning

Jianzhong You, Yuan Gao, Sangwook Kim, Chris Mcintosh

机构 * Peter Munk Cardiac Centre(彼得·芒克心脏中心) University Health Network (UHN)(大学健康网络) University of Toronto (U of T)(多伦多大学) Ted Rogers Centre for Heart Research(泰德·罗杰斯心脏研究中心) Toronto General Hospital Research Institute(多伦多综合医院研究所) Vector Institute(矢量研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07306 2025-03-11 cs.CL 57%

Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies

Luyi Jiang, Jiayuan Chen, Lu Lu, Xinwei Peng, Lihao Liu, Junjun He, Jie Xu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05204 2025-03-10 cs.CV 57%

Data-Efficient Generalization for Zero-shot Composed Image Retrieval

Zining Chen, Zhicheng Zhao, Fei Su, Xiaoqin Zhang, Shijian Lu

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19202 2025-03-10 cs.CL 57%

LiGT: Layout-infused Generative Transformer for Visual Question Answering on Vietnamese Receipts

Thanh-Phong Le, Trung Le Chi Phan, Nghia Hieu Nguyen, Kiet Van Nguyen

机构 * University of Information Technology(信息科技大学) Vietnam National University(越南国立大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted at IJDAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07335 2025-03-10 cs.LG cs.AI 57%

TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding

Haochuan Zhang, Chunhua Yang, Jie Han, Liyang Qin, Xiaoli Wang

机构 * Central South University(中南大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07527 2025-03-06 cs.CL 57%

Prompt-enhanced Network for Hateful Meme Classification

Junxi Liu, Yanyan Feng, Jiehai Chen, Yun Xue, Fenghuan Li

机构 * School of Electronics and Information Engineering, South China Normal University(华南师范大学电子与信息工程学院) School of Computer Science and Technology, Guangdong University of Technology(广东工业大学计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Published in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence Main Track. Pages 6397-6405

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00528 2025-03-04 cs.LG cs.CV 57%

Efficient Prompting for Continual Adaptation to Missing Modalities

Zirun Guo, Shulei Wang, Wang Lin, Weicai Yan, Yangyang Wu, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted to NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17000 2025-02-25 cs.CV cs.LG 57%

An Enhanced Large Language Model For Cross Modal Query Understanding System Using DL-KeyBERT Based CAZSSCL-MPGPT

Shreya Singh

机构 * DIT University(DIT大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 26 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13576 2025-02-25 cs.CL cs.IR 57%

FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research

Jiajie Jin, Yutao Zhu, Guanting Dong, Yuyao Zhang, Xinyu Yang, Chenghao Zhang, Tong Zhao, Zhao Yang, Zhicheng Dou, Ji-Rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments The paper is accepted by WWW2025 Resource Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13769 2025-02-20 cs.AI 57%

A consensus set for the aggregation of partial rankings: the case of the Optimal Set of Bucket Orders Problem

Juan A. Aledo, José A. Gámez, Alejandro Rosete

机构 * Universidad de Castilla-La Mancha(卡斯蒂利亚-拉曼恰大学) Universidad Tecnológica de La Habana Jose Antonio Echeverría(何塞·安东尼奥·埃切维里亚哈瓦那理工大学) Avangenio S.R.L.(阿万赫尼奥有限公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 26 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏