arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2306.06048 2024-07-30 cs.CV cs.CY cs.LG 57%

How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?

Yifei Ming, Yixuan Li

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IJCV 2023

Journal ref International Journal of Computer Vision 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18121 2024-07-26 cs.CV 57%

Efficient Inference of Vision Instruction-Following Models with Elastic Cache

Zuyan Liu, Benlin Liu, Jiahui Wang, Yuhao Dong, Guangyi Chen, Yongming Rao, Ranjay Krishna, Jiwen Lu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13129 2024-07-26 cs.CV cs.RO 57%

Better Call SAL: Towards Learning to Segment Anything in Lidar

Aljoša Ošep, Tim Meinhardt, Francesco Ferroni, Neehar Peri, Deva Ramanan, Laura Leal-Taixé

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14563 2024-07-23 cs.CV 57%

Learning Visual Grounding from Generative Vision and Language Model

Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun, Weicheng Kuo

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13808 2024-07-22 cs.CV 57%

CoAPT: Context Attribute words for Prompt Tuning

Gun Lee, Subin An, Sungyong Baik, Soochahn Lee

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12442 2024-07-18 cs.CV 57%

ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference

Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ECCV 2024. code available at https://github.com/mc- lan/ClearCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17507 2024-07-17 cs.CV 57%

HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts

Wonjae Kim, Sanghyuk Chun, Taekyung Kim, Dongyoon Han, Sangdoo Yun

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments ECCV 2024; 33pages, 4.5MB

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11001 2024-07-17 cs.CL cs.LG 57%

Generative AI Systems: A Systems-based Perspective on Generative AI

Jakub M. Tomczak

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16497 2024-07-16 cs.CV cs.LG 57%

PathoTune: Adapting Visual Foundation Model to Pathological Specialists

Jiaxuan Lu, Fang Yan, Xiaofan Zhang, Yue Gao, Shaoting Zhang

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08811 2024-07-15 eess.IV cs.CV 57%

CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting

Naman Sharma

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Supervised by Professor Ben Glocker

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09377 2024-07-15 cs.CV 57%

Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks

Tingyu Qu, Tinne Tuytelaars, Marie-Francine Moens

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08156 2024-07-12 cs.CV 57%

AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization

Shixiong Xu, Chenghao Zhang, Lubin Fan, Gaofeng Meng, Shiming Xiang, Jieping Ye

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03623 2024-07-12 cs.CV 57%

Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes

Yusuke Hirota, Jerone T. A. Andrews, Dora Zhao, Orestis Papakyriakopoulos, Apostolos Modas, Yuta Nakashima, Alice Xiang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00927 2024-07-12 cs.LG cs.AI stat.ML 57%

Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP

Zixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan Gu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 31 pages, 7 tables, 6 figures. In ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07360 2024-07-11 cs.CV cs.LG 57%

Towards a text-based quantitative and explainable histopathology image analysis

Anh Tien Nguyen, Trinh Thi Le Vuong, Jin Tae Kwak

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments MICCAI 2024 - Early acceptance (Top 11%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07341 2024-07-10 cs.LG cs.CR cs.CV 57%

Does CLIP Know My Face?

Dominik Hintersdorf, Lukas Struppek, Manuel Brack, Felix Friedrich, Patrick Schramowski, Kristian Kersting

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Published in the Journal of Artificial Intelligence Research (JAIR)

Journal ref Journal of Artificial Intelligence Research (JAIR), Vol. 80 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19092 2024-07-09 cs.CV 57%

Benchmarking and Improving Detail Image Caption

Hongyuan Dong, Jiawen Li, Bohong Wu, Jiacong Wang, Yuan Zhang, Haoyuan Guo

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04212 2024-07-08 cs.AI 57%

Smart Vision-Language Reasoners

Denisa Roberts, Lucas Roberts

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted in ICML 2024 MATH AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04068 2024-07-08 cs.CV 57%

CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-aware Prompting

Qinkai Yu, Jianyang Xie, Anh Nguyen, He Zhao, Jiong Zhang, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03380 2024-07-08 q-bio.QM cs.AI cs.LG 57%

Multi-Peptide: Multimodality Leveraged Language-Graph Learning of Peptide Properties

Srivathsan Badrinarayanan, Chakradhar Guntuboina, Parisa Mollaei, Amir Barati Farimani

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06879 2024-07-08 cs.CV eess.IV 57%

The Solution for the CVPR2023 NICE Image Captioning Challenge

Xiangyu Wu, Yi Gao, Hailiang Zhang, Yang Yang, Weili Guo, Jianfeng Lu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12752 2024-07-03 cs.CV 57%

C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning

Ji Ma, Wei Suo, Peng Wang, Yanning Zhang

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI-24

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00905 2024-07-02 cs.CV 57%

Learning Robust 3D Representation from CLIP via Dual Denoising

Shuqing Luo, Bowen Qu, Wei Gao

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09233 2024-07-02 cs.CV 57%

Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP

Samyadeep Basu, Shell Xu Hu, Maziar Sanjabi, Daniela Massiceti, Soheil Feizi

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00556 2024-07-02 cs.MM 57%

Revisiting Vision-Language Features Adaptation and Inconsistency for Social Media Popularity Prediction

Chih-Chung Hsu, Chia-Ming Lee, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Yu Jian, Chi-Han Tsai

专题命中 图文多模态 :multi-modal(abstract);分类 cs.MM

Comments Submission of the 7th Social Media Prediction Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00252 2024-07-02 cs.CV cs.ET 57%

Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review

Moseli Mots'oehli

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted IEEE ETNCC 2024, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03118 2024-06-26 cs.CV 57%

LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Gabriela Ben Melech Stan, Estelle Aflalo, Raanan Yehezkel Rohekar, Anahita Bhiwandiwalla, Shao-Yen Tseng, Matthew Lyle Olson, Yaniv Gurwicz, Chenfei Wu, Nan Duan, Vasudev Lal

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00789 2024-06-21 cs.CV 57%

Retrieval-Augmented Egocentric Video Captioning

Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, Weidi Xie

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2024. Project page is available at: https://jazzcharles.github.io/Egoinstructor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12668 2024-06-19 cs.CV 57%

Disturbing Image Detection Using LMM-Elicited Emotion Embeddings

Maria Tzelepi, Vasileios Mezaris

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication, LVLM Workshop @ IEEE Int. Conf. on Image Processing (ICIP 2024), Abu Dhabi, United Arab Emirates, Oct. 2024. This is the authors' "accepted version"

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12225 2024-06-19 cs.CV 57%

The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge

Hongpeng Pan, Shifeng Yi, Shouwei Yang, Lei Qi, Bing Hu, Yi Xu, Yang Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR2024 Foundational Few-Shot Object Detection Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏