arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2504.20468 2025-05-08 cs.CV 57%

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception

Yuanchen Wu, Lu Zhang, Hang Yao, Junlong Du, Ke Yan, Shouhong Ding, Yunsheng Wu, Xiaoqiang Li

机构 * School of Computer Engineering & Science, Shanghai University(上海大学计算机工程与科学学院) Tencent YouTu Lab(腾讯优图实验室)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07971 2025-05-08 cs.CV 57%

XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Omkar Thawakar, Abdelrahman Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Aalto University(艾尔沃斯大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted at ACL 2024-BIONLP Workshop. Code: https://github.com/mbzuai-oryx/XrayGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03350 2025-05-07 cs.CV 57%

A Vision-Language Model for Focal Liver Lesion Classification

Song Jian, Hu Yuchang, Wang Hui, Chen Yen-Wei

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages,4 figures, 4 tables,Innovation in Medicine and Healthcare Proceedings of 13th KES-InMed 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15326 2025-05-07 cs.CV 57%

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

Jiajie Li, Brian R Quaranto, Chenhui Xu, Ishan Mishra, Ruiyang Qin, Dancheng Liu, Peter C W Kim, Jinjun Xiong

机构 * Department of Computer Science and Engineering, University at Buffalo(计算机科学与工程系,布法罗大学) Department of Surgery, University at Buffalo(外科系,布法罗大学) Department of Computer Science and Engineering, IIT, Jodhpur(计算机科学与工程系,印度理工学院,乔杜尔) Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02829 2025-05-06 cs.AI 57%

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta, Trevor Darrell, David M. Chan

机构 * Department of Electrical Engineering and Computer Sciences, University of California-Berkeley, Berkeley, CA, USA(电气工程与计算机科学系,加州大学伯克利分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 28 pages, 10 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02182 2025-05-06 cs.CV 57%

Robust AI-Generated Face Detection with Imbalanced Data

Yamini Sri Krubha, Aryana Hou, Braden Vester, Web Walker, Xin Wang, Li Lin, Shu Hu

机构 * Purdue University(普渡大学) Clarkstown High School South(Clarkstown 高中南分校) University at Albany, State University of New York(阿尔巴尼大学,纽约州立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14467 2025-05-02 cs.CV 57%

LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation

Jiachen Li, Qing Xie, Renshu Gu, Jinyu Xu, Yongjian Liu, Xiaohan Yu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11672 2025-05-02 cs.AI cs.LG 57%

Artificial Scientific Discovery

Antonio Norelli

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments PhD thesis, 123 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12286 2025-05-02 cs.RO cs.CV 57%

GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Teli Ma, Zifan Wang, Jiaming Zhou, Mengmeng Wang, Junwei Liang

机构 * AI, HKUST(GZ)(香港科技大学(广州)人工智能学院) ZJUT(浙江工业大学) CSE, HKUST(香港科技大学计算机科学与工程系)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20152 2025-05-01 cs.CV 57%

Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals

Phillip Howard, Kathleen C. Fraser, Anahita Bhiwandiwalla, Svetlana Kiritchenko

机构 * Intel Labs(英特尔实验室) National Research Council Canada(加拿大国家研究理事会)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NAACL 2025 main track (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13181 2025-04-30 cs.CV 57%

Perception Encoder: The best visual embeddings are not at the output of the network

Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei, Tengyu Ma, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Junke Wang, Marco Monteiro, Hu Xu, Shiyu Dong, Nikhila Ravi, Daniel Li, Piotr Dollár, Christoph Feichtenhofer

机构 * Meta FAIR Fudan University(复旦大学) Meta Reality Labs

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Updated refs, fixed typos, and added new COCO SotA: 66.0 val mAP! Code, models, and data at https://github.com/facebookresearch/perception_models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19524 2025-04-29 cs.CV 57%

LR-IAD:Mask-Free Industrial Anomaly Detection with Logical Reasoning

Peijian Zeng, Feiyan Pang, Zhanbo Wang, Aimin Yang

机构 * School of Computer Science and Technology, Guangdong University of Technology(广东技术大学计算机科学与技术学院) School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19086 2025-04-29 cs.CV 57%

Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction

Xiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi, Shichu Sun, Jing Xing, Jian Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(先进交叉学科研究院,中国科学院大学) Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18738 2025-04-29 cs.CV 57%

A Review of 3D Object Detection with Vision-Language Models

Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally, Marco Flores Calero, Manoj Karkee

机构 * Cornell University(康奈尔大学) University of Peloponnese(希腊皮洛斯大学) Kansas State University(堪萨斯州立大学) Universidad de las Fuerzas Armadas(武装力量大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07149 2025-04-29 cs.CV cs.LG 57%

Towards Interpreting Visual Information Processing in Vision-Language Models

Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez

机构 * Nanyang Technological University(南洋理工大学) University of Oxford(牛津大学) Tel Aviv University(特拉维夫大学) MILA(蒙特利尔人工智能研究院) ERA-Krueger AI Safety Lab(ERA-Krueger人工智能安全实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Published at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18125 2025-04-29 cs.CV 57%

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Project page: https://zcmax.github.io/projects/LLaVA-3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05091 2025-04-28 cs.CV 57%

DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions

Gorkem Can Ates, Yu Xin, Kuang Gong, Wei Shao

机构 * Department of Medicine, University of Florida(医学系,佛罗里达大学) Department of Electrical and Computer Engineering, University of Florida(电气与计算机工程系,佛罗里达大学) Department of Biomedical Engineering, University of Florida(生物医学工程系,佛罗里达大学) Intelligent Clinical Care Center, University of Florida(智能临床护理中心,佛罗里达大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15271 2025-04-22 cs.CV 57%

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tuomas Rintamaki, Tyler Poon, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14692 2025-04-22 cs.CL 57%

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

Songtao Jiang, Yuan Wang, Sibo Song, Yan Zhang, Zijie Meng, Bohan Lei, Jian Wu, Jimeng Sun, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) UIUC(伊利诺伊大学香槟分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12206 2025-04-22 cs.CV 57%

TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification

Ans Munir, Faisal Z. Qureshi, Muhammad Haris Khan, Mohsen Ali

机构 * Information Technology University(信息科技大学) University of Ontario Institute of Technology(Ontario Institute of Technology 大学) Mohamed Bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Added code link in the abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11733 2025-04-22 cs.CV 57%

DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment

Li Yu, Situo Wang, Wei Zhou, Moncef Gabbouj

机构 * School of Computer Science, Nanjing University of Information Science & Technology(信息科学技术南京大学计算机科学学院) Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET)(江苏省大气环境与装备技术协同创新中心) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学) Faculty of Information Technology and Communication Sciences, Tampere University(信息科技与通信科学学院,坦佩雷大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11468 2025-04-17 cs.CL 57%

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Hardy Chen, Haoqin Tu, Fali Wang, Hui Liu, Xianfeng Tang, Xinya Du, Yuyin Zhou, Cihang Xie

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11111 2025-04-16 cs.CV 57%

S$^2$Teacher: Step-by-step Teacher for Sparsely Annotated Oriented Object Detection

Yu Lin, Jianghang Lin, Kai Ye, You Shen, Yan Zhang, Shengchuan Zhang, Liujuan Cao, Rongrong Ji

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10320 2025-04-15 cs.CV 57%

SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model

Zongcan Ding, Haodong Zhang, Peng Wu, Guansong Pang, Zhiwei Yang, Peng Wang, Yanning Zhang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11609 2025-04-15 cs.CV 57%

CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution

Bingwen Hu, Heng Liu, Zhedong Zheng, Ping Liu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08140 2025-04-14 cs.CV 57%

Impact of Language Guidance: A Reproducibility Study

Cherish Puniani, Advika Sinha, Shree Singhi, Aayan Yadav

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07165 2025-04-11 cs.CV 57%

Perception in Reflection

Yana Wei, Liang Zhao, Kangheng Lin, En Yu, Yuang Peng, Runpei Dong, Jianjian Sun, Haoran Wei, Zheng Ge, Xiangyu Zhang, Vishal M. Patel

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05227 2025-04-08 cs.CV 57%

A Reality Check of Vision-Language Pre-training in Radiology: Have We Progressed Using Text?

Julio Silva-Rodríguez, Jose Dolz, Ismail Ben Ayed

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments IPMI 2025. Code and weights: https://github.com/jusiro/DLILP

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04470 2025-04-08 cs.CV 57%

Domain Generalization for Face Anti-spoofing via Content-aware Composite Prompt Engineering

Jiabao Guo, Ajian Liu, Yunfeng Diao, Jin Zhang, Hui Ma, Bo Zhao, Richang Hong, Meng Wang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02862 2025-04-08 cs.CV cs.LG 57%

Towards Understanding How Knowledge Evolves in Large Vision-Language Models

Sudong Wang, Yunjian Zhang, Yao Zhu, Jianing Li, Zizhe Wang, Yanwei Liu, Xiangyang Ji

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏