arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-31 至 2025-07-31 共收录 41 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2507.22431 2025-07-31 cs.CV 83%

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

Zhixiang Wei, Guangting Wang, Xiaoxiao Ma, Ke Mei, Huaian Chen, Yi Jin, Fengyun Rao

机构 * University of Science and Technology of China(中国科学技术大学) WeChat Vision, Tencent Inc.(腾讯公司)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22074 2025-07-31 cs.LG cs.CL 83%

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

Yangshu Yuan, Heng Chen, Xinyi Jiang, Christian Ng, Kexin Qiu

机构 * Singapore Institute of Management(新加坡管理学院)

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04585 2025-07-31 cs.CL 79%

MuSciClaims: Multimodal Scientific Claim Verification

Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi, Mahnaz Koupaee, Niranjan Balasubramanian

机构 * Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14200 2025-07-31 cs.CV cs.AI 76%

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization

Huiyi Chen, Jiawei Peng, Kaihua Tang, Xin Geng, Xu Yang

机构 * Southeast University(东南大学) Huawei Singapore Research Center(华为新加坡研究中心)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22346 2025-07-31 cs.CV 57%

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

Pei Deng, Wenqian Zhou, Hanlin Wu

机构 * School of Information Science and Technology, Beijing Foreign Studies University(信息科学与技术学院,北京外国语大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 5 figures. Submitted to IEEE Transactions on Geoscience and Remote Sensing (TGRS). Code and dataset are available at https://github.com/hanlinwu/DeltaVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22003 2025-07-31 cs.CV 57%

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

Ziyun Dai, Xiaoqiang Li, Shaohua Zhang, Yuanchen Wu, Jide Li

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM25

Journal ref 33rd ACM International Conference on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12447 2025-07-31 cs.CV 57%

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification

Nathanael L. Baisa, Babu Pallam, Amudhavel Jayavel

机构 * School of Computer Science(计算机科学学院) Informatics De Montfort University(信息学德蒙特福特大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22101 2025-07-31 cs.CV 57%

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock

Umair Nawaz, Muhammad Zaigham Zaheer, Fahad Shahbaz Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer

机构 * MBZ University of AI(人工智能大学) CECS, Australian National University(计算机科学与工程系,澳大利亚国立大学) Computer Vision Laboratory, Linköping University(链接öping大学计算机视觉实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 图文多模态 :multimodal(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2507.22781 2025-07-31 cs.CV 85%

HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training

Xuecheng Wu, Danlei Huang, Heli Sun, Xinyi Yin, Yifan Wang, Hao Wang, Jia Zhang, Fei Wang, Peihao Guo, Suyu Xing, Junxiao Xue, Liang He

机构 * Xi'an Jiaotong University(西安交通大学) Zhengzhou University(郑州大学) University of Science and Technology of China(中国科学技术大学) Dalian University of Technology(大连理工大学) Zhejiang Lab(浙江实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04906 2025-07-31 cs.MM cs.CV cs.SD eess.AS 78%

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio

机构 * Department of Computer Science, University of Bari Aldo Moro, Italy(巴里阿尔多·莫罗大学计算机科学系)

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV、cs.MM、eess.AS

Comments Presented at the AI for Visual Arts (AI4VA) workshop at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11026 2025-07-31 eess.AS cs.CV cs.LG cs.MM 75%

MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

Sungwoo Cho, Jeongsoo Choi, Sungnyun Kim, Se-Young Yun

机构 * KAIST AI(韩国科学技术院人工智能研究所) KAIST EE(韩国科学技术院电子工程系)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20999 2025-07-31 cs.SD cs.GR eess.AS 57%

Text-Driven Voice Conversion via Latent State-Space Modeling

Wen Li, Sofia Martinez, Priyanka Shah

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11308 2025-07-31 eess.AS eess.SP 57%

Uncovering the role of semantic and acoustic cues in normal and dichotic listening

Sai Samrat Kankanala, Akshara Soman, Sriram Ganapathy

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 10 Pages, 6 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2411.07076 2025-07-31 cs.CV cs.AI 84%

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang, Yuchen Zhang, Ruicheng Le

专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20988 2025-07-31 cs.CL cs.GR 83%

Cross-Modal State-Space Graph Reasoning for Structured Summarization

Hannah Kim, Sofia Martinez, Jason Lee

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22878 2025-07-31 cs.IR cs.CL cs.CY 79%

GeoOutageKG: A Multimodal Geospatiotemporal Knowledge Graph for Multiresolution Power Outage Analysis

Ethan Frakes, Yinghui Wu, Roger H. French, Mengjie Li

机构 * University of Central Florida(佛罗里达中央大学) Case Western Reserve University(凯斯西储大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to the 24th International Semantic Web Conference Resource Track (ISWC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22229 2025-07-31 cs.LG 50%

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville, Jean-Rémi King

机构 * Meta AI

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2503.16133 2025-07-31 cs.GR 50%

Multi-Prompt Style Interpolation for Fine-Grained Artistic Control

Lei Chen, Hao Li, Yuxin Zhang, Chao Li, Kai Wen

专题命中 跨模态检索 :cross-modal(abstract)

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2507.21391 2025-07-31 cs.CV cs.AI cs.CL 85%

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

Shijie Zhou, Ruiyi Zhang, Huaisheng Zhu, Branislav Kveton, Yufan Zhou, Jiuxiang Gu, Jian Chen, Changyou Chen

机构 * University at Buffalo(布法罗大学) Adobe Research(Adobe研究) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at ICCV 2025. Code available at https://github.com/sjz5202/LLaVA-Reward

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20132 2025-07-31 cs.LG 78%

High-Resolution Live Fuel Moisture Content (LFMC) Maps for Wildfire Risk from Multimodal Earth Observation Data

Patrick Alan Johnson, Gabriel Tseng, Yawen Zhang, Heather Heward, Virginia Sjahli, Favyen Bastani, Joseph Redmon, Patrick Beukema

机构 * Mila -- Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Allen Institute for AI (Ai2)(人工智能研究所) University of Idaho(爱达荷大学)

专题命中 多模态生成 :multimodal(title,abstract)

Comments 10 pages, ICML 2025 (TerraBytes)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00983 2025-07-31 eess.IV cs.CV 70%

DMCIE: Diffusion Model with Concatenation of Inputs and Errors to Improve the Accuracy of the Segmentation of Brain Tumors in MRI Images

Sara Yavari, Rahul Nitin Pandya, Jacob Furst

机构 * School of Computing, DePaul University(计算学院,德保罗大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22076 2025-07-31 cs.LG 67%

Test-time Prompt Refinement for Text-to-Image Models

Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya, Vibhav Vineet

机构 * Florida Institute of Technology(佛罗里达理工学院) Microsoft Research(微软研究院)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract)

Comments Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2507.22367 2025-07-31 cs.CL cs.MM 84%

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors

Jia Li, Yichao He, Jiacheng Xu, Tianhao Luo, Zhenzhen Hu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学)

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);audio-visual(abstract);分类 cs.CL、cs.MM

Comments 8 pages, 3 figures, ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22685 2025-07-31 cs.CV cs.AI 81%

Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing

Yimeng Liu, Maolin Gan, Yidong Ren, Gen Li, Jingkai Lin, Younsuk Dong, Zhichao Cao

机构 * Michigan State University(密歇根州立大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22412 2025-07-31 cs.CV 79%

UAVScenes: A Multi-Modal Dataset for UAVs

Sijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu, Shenghai Yuan, Rui She, Quanjiang Guo, JinXuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas, Toshan Aggarwal, Heyuan Liu, Hongming Zhang, Chujie Chen, Junyu Jiang, Lihua Xie, Wee Peng Tay

机构 * Nanyang Technological University(南洋理工大学) School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Beihang University(北航大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22099 2025-07-31 cs.CV cs.AI cs.MM cs.SE 67%

Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?

Shuqing Li, Qiang Chen, Xiaoxue Ren, Michael R. Lyu

机构 * The Chinese University of Hong Kong, China(香港中文大学) State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19891 2025-07-31 cs.CV cs.AI 62%

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention

Drandreb Earl O. Juanico, Rowel O. Atienza, Jeffrey Kenneth Go

机构 * AI Graduate Program, University of the Philippines(菲律宾大学人工智能研究生项目) EEEI, University of the Philippines(菲律宾大学电子工程与信息学院) Samsung R&D Institute Philippines(三星菲律宾研发研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments To be published in the ICCVW 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09950 2025-07-31 cs.CV cs.AI 62%

Can GPT-4o mini and Gemini 2.0 Flash Predict Fine-Grained Fashion Product Attributes? A Zero-Shot Analysis

Shubham Shukla, Kunal Sonalkar

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Version 2: Added a missing citation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10696 2025-07-31 cs.RO cs.CV 57%

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

Manthan Patel, Fan Yang, Yuheng Qiu, Cesar Cadena, Sebastian Scherer, Marco Hutter, Wenshan Wang

机构 * Robotic Systems Lab, ETH Zurich(苏黎世联邦理工学院机器人系统实验室) Robotics Institute of Carnegie Mellon University(卡内基梅隆大学机器人研究所)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted for publication to IEEE/RSJ IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏