arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2306.02092 2024-09-04 cs.CV 57%

Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations

Xu Zhang, Zhedong Zheng, Linchao Zhu, Yi Yang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted by Knowledge-Based Systems (KBS)

Journal ref Knowl. Based. Syst. 300 (2024) 112135

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06762 2024-08-29 cs.AI 57%

Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions

Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments ECAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12981 2024-08-26 cs.AI 57%

QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval

Chenghua Gao, Min Li, Jianshuo Liu, Junxing Ren, Lin Chen, Haoyu Liu, Bo Meng, Jitao Fu, Wenwen Su

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

Comments 9 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08575 2024-08-19 cs.CV 57%

Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs

Jinming Liu, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen, Wenjun Zeng, Xin Jin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03586 2024-08-15 cs.CV eess.IV 57%

SSL-SoilNet: A Hybrid Transformer-based Framework with Self-Supervised Learning for Large-scale Soil Organic Carbon Prediction

Nafiseh Kakhani, Moien Rangzan, Ali Jamali, Sara Attarchi, Seyed Kazem Alavipanah, Michael Mommert, Nikolaos Tziolas, Thomas Scholten

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in IEEE Transactions on Geoscience and Remote Sensing (TGRS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02392 2024-08-06 cs.CV eess.IV 57%

MaFreeI2P: A Matching-Free Image-to-Point Cloud Registration Paradigm with Active Camera Pose Retrieval

Gongxin Yao, Xinyang Li, Yixin Xuan, Yu Pan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Conference on Multimedia Expo 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20243 2024-07-31 cs.CL cs.LG 57%

Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions

Jinsung Yoon, Raj Sinha, Sercan O Arik, Tomas Pfister

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17671 2024-07-29 cs.CV cs.LG 57%

Unsqueeze [CLS] Bottleneck to Learn Rich Representations

Qing Su, Shihao Ji

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12713 2024-07-23 cs.CV 57%

Dynamic Identity-Guided Attention Network for Visible-Infrared Person Re-identification

Peng Gao, Yujian Lee, Hui Zhang, Xubo Liu, Yiyang Hu, Guquan Jing

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments I need to further debug my code to improve accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14026 2024-07-22 cs.CV 57%

Semi-supervised reference-based sketch extraction using a contrastive learning framework

Chang Wook Seo, Amirsaman Ashtari, Junyong Noh

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Main paper 1-12 page, Supplementary 13-34 page

Journal ref ACM Transactions on Graphics (TOG) 2023, Volume 42, Issue 4 Article No.: 56, Pages 1 - 12

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13773 2024-07-22 cs.DL cs.AI 57%

OpenDataLab: Empowering General Artificial Intelligence with Open Datasets

Conghui He, Wei Li, Zhenjiang Jin, Chao Xu, Bin Wang, Dahua Lin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08530 2024-07-18 cs.CV 57%

X-Pose: Detecting Any Keypoints

Jie Yang, Ailing Zeng, Ruimao Zhang, Lei Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments ECCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11816 2024-07-16 cs.CV cs.LG 57%

Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning

Jihai Zhang, Xiang Lan, Xiaoye Qu, Yu Cheng, Mengling Feng, Bryan Hooi

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments ECCV 2024 Camera-Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05898 2024-07-09 cs.AI 57%

Contrastive Learning of Preferences with a Contextual InfoNCE Loss

Timo Bertram, Johannes Fürnkranz, Martin Müller

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03615 2024-07-08 cs.CL 57%

Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models

Chang-Sheng Kao, Yun-Nung Chen

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15935 2024-07-04 cs.LG cs.AI 57%

MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences

Viktor Moskvoretskii, Dmitry Osin, Egor Shvetsov, Igor Udovichenko, Maxim Zhelnin, Andrey Dukhovny, Anna Zhimerikina, Evgeny Burnaev

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01810 2024-07-03 cs.CV 57%

Freeview Sketching: View-Aware Fine-Grained Sketch-Based Image Retrieval

Aneeshan Sain, Pinaki Nath Chowdhury, Subhadeep Koley, Ayan Kumar Bhunia, Yi-Zhe Song

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted in European Conference on Computer Vision (ECCV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05646 2024-07-03 cs.CV 57%

Masked Attribute Description Embedding for Cloth-Changing Person Re-identification

Chunlei Peng, Boyu Wang, Decheng Liu, Nannan Wang, Ruimin Hu, Xinbo Gao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01219 2024-07-02 cs.CL 57%

Searching for Best Practices in Retrieval-Augmented Generation

Xiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, Ruicheng Yin, Changze Lv, Xiaoqing Zheng, Xuanjing Huang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18400 2024-06-28 cs.MM 57%

Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning

Hanyao Wang, Yibing Zhan, Liu Liu, Liang Ding, Yan Yang, Jun Yu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM

Comments This work has been submitted to the lEEE for possible publication. Copyright may betransferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12881 2024-06-21 physics.acc-ph cs.CL 57%

Towards Unlocking Insights from Logbooks Using AI

Antonin Sulc, Alex Bien, Annika Eichler, Daniel Ratner, Florian Rehm, Frank Mayet, Gregor Hartmann, Hayden Hoschouer, Henrik Tuennermann, Jan Kaiser, Jason St. John, Jennefer Maldonado, Kyle Hazelwood, Raimund Kammering, Thorsten Hellert, Tim Wilksen, Verena Kain, Wan-Lin Hu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments 5 pages, 1 figure, 15th International Particle Accelerator Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05620 2024-06-11 cs.CV 57%

Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval

Yiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang, Weilin Zhuang, Rongrong Ji

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments ACM MM2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05344 2024-06-11 cs.CL 57%

MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention

Prince Jha, Raghav Jain, Konika Mandal, Aman Chadha, Sriparna Saha, Pushpak Bhattacharyya

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07104 2024-06-07 cs.AI cs.PL 57%

SGLang: Efficient Execution of Structured Language Model Programs

Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05001 2024-06-07 cs.CV cs.LG 57%

COURIER: Contrastive User Intention Reconstruction for Large-Scale Visual Recommendation

Jia-Qi Yang, Chenglei Dai, Dan OU, Dongshuai Li, Ju Huang, De-Chuan Zhan, Xiaoyi Zeng, Yang Yang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05634 2024-06-04 cs.CV 57%

PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification

Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, Ngan Le, Minh-Triet Tran

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at AVSS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05510 2024-05-29 cs.CV 57%

Mutimodal Ranking Optimization for Heterogeneous Face Re-identification

Hui Hu, Jiawei Zhang, Zhen Han

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments In the methods section, unpaired face samples should be used for training during negative sample training, rather than paired samples. Corresponding errors are also present in the subsequent experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16328 2024-05-28 cs.CV 57%

A Classifier-Free Incremental Learning Framework for Scalable Medical Image Segmentation

Xiaoyang Chen, Hao Zheng, Yifang Xie, Yuncong Ma, Tengfei Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02141 2024-05-20 cs.CV 57%

Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization

Bo Yang, Chen Wang, Xiaoshuang Ma, Beiping Song, Zhuang Liu, Fangde Sun

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 44 pages, 6 figures

Journal ref Remote Sens. 2024, 16, 1653

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09594 2024-05-17 eess.IV cs.CV cs.LG 57%

Learning Generalized Medical Image Representations through Image-Graph Contrastive Pretraining

Sameer Khanna, Daniel Michael, Marinka Zitnik, Pranav Rajpurkar

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted into Machine Learning for Health (ML4H) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏