arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-28 至 2025-07-28 共收录 36 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2507.19370 2025-07-28 cs.CV 74%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18915 2025-07-28 cs.CL cs.CV 62%

Mining Contextualized Visual Associations from Images for Creativity Understanding

Ananya Sahu, Amith Ananthram, Kathleen McKeown

机构 * Columbia University(哥伦比亚大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05211 2025-07-28 cs.CV cs.AI 62%

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫兹哈德大学人工智能大学) Technical University of Munich(慕尼黑技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15265 2025-07-28 cs.CV cs.AI cs.CR 62%

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

Zihao Pan, Yu Tong, Weibin Wu, Jingyi Wang, Lifeng Chen, Zhe Zhao, Jiajia Wei, Yitong Qiao, Zibin Zheng

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper needs major revisions, so it is being withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2507.19356 2025-07-28 cs.CL cs.SD eess.AS 62%

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen

机构 * Department of English(英语系) National Taiwan Normal University(台湾师范大学) Department of Computer Science and Information Engineering(计算机科学与信息工程系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments 6 pages, 3 figures, to appear in the Proceedings of the 2025 International Conference on Asian Language Processing (IALP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18380 2025-07-28 cs.AI cs.LG 57%

RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification

Praphul Singh, Charlotte Dzialo, Jangwon Kim, Sumana Srivatsa, Irfan Bulu, Sri Gadde, Krishnaram Kenthapadi

机构 * Oracle Health & AI(Oracle健康与人工智能)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted to ACL 2025 Industry Track. To appear

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2507.18104 2025-07-28 cs.CV q-bio.NC 79%

A Multimodal Seq2Seq Transformer for Predicting Brain Responses to Naturalistic Stimuli

Qianyi He, Yuan Chang Leong

机构 * Data Science Institute University of Chicago(芝加哥大学数据科学研究所) Department of Psychology, Neuroscience Institute University of Chicago(芝加哥大学心理学系、神经科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12561 2025-07-28 cs.CV cs.AI 73%

Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

Wenjun Huang, Yang Ni, Hanning Chen, Yirui He, Ian Bryant, Yezi Liu, Mohsen Imani

机构 * University of California, Irvine(加州大学尔湾分校)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17958 2025-07-28 cs.LG cs.AI cs.CV 62%

VIBE: Video-Input Brain Encoder for fMRI Response Modeling

Daniel Carlström Schad, Shrey Dixit, Janis Keck, Viktor Studenyak, Aleksandr Shpilevoi, Andrej Bicanski

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19082 2025-07-28 cs.RO 50%

Bot Appétit! Exploring how Robot Morphology Shapes Perceived Affordances via a Mise en Place Scenario in a VR Kitchen

Rachel Ringe, Leandra Thiele, Mihai Pomarlan, Nima Zargham, Robin Nolte, Lars Hurrelbrink, Rainer Malaka

机构 * Digital Media Lab, University of Bremen(柏林布雷门大学数字媒体实验室) Department of Linguistics, University of Bremen(柏林布雷门大学语言学系)

专题命中 视频多模态 :multimodal(abstract)

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2507.19054 2025-07-28 cs.CV cs.AI cs.CL cs.IR cs.LG 67%

Closing the Modality Gap for Mixed Modality Search

Binxu Li, Yuhui Zhang, Xiaohan Wang, Weixin Liang, Ludwig Schmidt, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page: https://yuhui-zh15.github.io/MixedModalitySearch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18881 2025-07-28 cs.CV cs.RO 57%

Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?

Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong, Jianxin Wang

机构 * School of Computer Science Engineering, Central South University Changsha Hunan China Engineering, Central South University

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12591 2025-07-28 cs.CL 57%

LLMs are Also Effective Embedding Models: An In-depth Overview

Chongyang Tao, Tao Shen, Shen Gao, Junshuo Zhang, Zhen Li, Kai Hua, Wenpeng Hu, Zhengwei Tao, Shuai Ma

机构 * Beihang University(北航大学) SKLSDE Lab, Beihang University(北航SKLSDE实验室) University of Technology Sydney(悉尼大学) University of Electronic Science and Technology of China(电子科技大学) Peking University(北京大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments 38 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2507.18763 2025-07-28 cs.CV cs.RO 79%

Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving

Keshav Gupta, Tejas S. Stanley, Pranjal Paul, Arun K. Singh, K. Madhava Krishna

机构 * Robotics Research Center, IIIT-Hyderabad(IIIT-海得拉巴机器人研究中心) The University of Tartu(塔尔图大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 7 figures, IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18750 2025-07-28 cs.MM cs.SD eess.AS 73%

CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

Hyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim

机构 * Hanyang University(翰阳大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07919 2025-07-28 cs.CL q-bio.BM 70%

Advancing biomolecular understanding and design following human instructions

Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, Qiang Zhang, Huajun Chen

专题命中 多模态生成 :multimodal(abstract);any-to-any(abstract);分类 cs.CL

Journal ref Nature Machine Intelligence volume 7, pages1154-1167 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18667 2025-07-28 cs.CV cs.AI 62%

Gen-AI Police Sketches with Stable Diffusion

Nicholas Fidalgo, Aaron Contreras, Katherine Harvey, Johnny Ni

机构 * Harvard College(哈佛学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 6 篇

2507.19213 2025-07-28 cs.CV 83%

PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction

Hanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao, Tianyu Fu, Minghui Wu, Beiping Tan, Huiying Li

机构 * Jilin University(吉林大学) Peking University(北京大学) Mininglamp Technology(Mininglamp科技)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19172 2025-07-28 cs.AI cs.CV 81%

PhysDrive: A Multimodal Remote Physiological Measurement Dataset for In-vehicle Driver Monitoring

Jiyao Wang, Xiao Yang, Qingyong Hu, Jiankai Tang, Can Liu, Dengbo He, Yuntao Wang, Yingcong Chen, Kaishun Wu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学) Sichuan Agricultural University(四川农业大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments It is the initial version, not the final version

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 79%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19304 2025-07-28 cs.CV cs.AI 62%

Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes

Muhammad Ibrahim, Naveed Akhtar, Haitian Wang, Saeed Anwar, Ajmal Mian

机构 * Department of Computer Science, The University of Western Australia(西澳大学计算机科学系) School of Computing & Information Systems, The University of Melbourne(墨尔本大学计算与信息系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This paper has been accepted by IEEE/RSJ IROS 2025 for oral presentation on 19 Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19083 2025-07-28 cs.CV cs.AI 62%

ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

Yuqian Fu, Runze Wang, Bin Ren, Guolei Sun, Biao Gong, Yanwei Fu, Danda Pani Paudel, Xuanjing Huang, Luc Van Gool

机构 * Fudan University(复旦大学) University of Trento(特伦托大学) University of Pisa(比萨大学) ETH Zürich(苏黎世联邦理工学院) Ant Group(蚂蚁集团)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV25 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11493 2025-07-28 cs.CV 57%

GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing

Yusu Qian, Jiasen Lu, Tsu-Jui Fu, Xinze Wang, Chen Chen, Yinfei Yang, Wenze Hu, Zhe Gan

机构 * Apple(苹果公司)

专题命中 多模态评测 :image-text(abstract);分类 cs.CV

Comments Project page: https://sueqian6.github.io/GIE-Bench-web/

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2507.19344 2025-07-28 math.OC 50%

Joint Inference of Trajectory and Obstacle in Mean-Field Games via Bilevel Optimization

Han Huang, Jiajia Yu, Tianyi Chen, Rongjie Lai

专题命中 多模态Agent :multi-modal(abstract)

Comments 17 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19096 2025-07-28 cs.NI 50%

iPLAN: Redefining Indoor Wireless Network Planning Through Large Language Models

Jinbo Hou, Stefanos Bakirtzis, Kehai Qiu, Sichong Liao, Hui Song, Haonan Hu, Kezhi Wang, Jie Zhang

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 5 篇

2411.19628 2025-07-28 cs.CV cs.CL cs.LG cs.MM 82%

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Qiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye, Zhanpeng Zen, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18940 2025-07-28 cs.CL cs.MM 81%

LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation

Jingxuan Wei, Caijun Jia, Qi Chen, Yujun Cai, Linzhuang Sun, Xiangxiang Zhang, Gaowei Wu, Bihui Yu

机构 * University of Chinese Academy of Sciences(中国科学院大学) The University of Queensland(昆士兰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18929 2025-07-28 cs.CV cs.AI 81%

MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition

Jian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang, Min Yang, Chengming Li, Xiping Hu

机构 * Shenzhen MSU-BIT University(深圳MSU-BIT大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19253 2025-07-28 cs.CV 79%

BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection

An Xiang, Zixuan Huang, Xitong Gao, Kejiang Ye, Cheng-zhong Xu

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Shenzhen University of Advanced Technology(深圳先进技术大学) State Key Lab of IOTSC, Department of CIS, University of Macau(物联网科学国家重点实验室,澳门大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19052 2025-07-28 cs.CV 79%

Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding

Hamid Abdollahi, Amir Hossein Mansouri Majoumerd, Amir Hossein Bagheri Baboukani, Amir Abolfazl Suratgar, Mohammad Bagher Menhaj

机构 * Distributed and Intelligent Optimization Research Laboratory(分布式智能优化研究实验室) Electrical Engineering Department(电气工程系) Amirkabir University of Technology(阿米尔卡比尔技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏