arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-24 至 2025-09-24 共收录 70 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2402.05930 2025-09-24 cs.CL cs.CV cs.LG 62%

WebLINX: Real-World Website Navigation with Multi-Turn Dialogue

Xing Han Lù, Zdeněk Kasner, Siva Reddy

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 5 篇

2506.06561 2025-09-24 cs.CL cs.AI cs.CV 82%

LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles

Ho Yin 'Sam' Ng, Ting-Yao Hsu, Aashish Anantha Ramakrishnan, Branislav Kveton, Nedim Lipka, Franck Dernoncourt, Dongwon Lee, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ting-Hao 'Kenneth' Huang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings. The LaMP-CAP dataset is publicly available at: https://github.com/Crowd-AI-Lab/lamp-cap

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18824 2025-09-24 cs.CV 79%

Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation

Yanzuo Lu, Xin Xia, Manlin Zhang, Huafeng Kuang, Jianbin Zheng, Yuxi Ren, Xuefeng Xiao

机构 * ByteDance Seed(字节跳动种子)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18461 2025-09-24 cs.GR cs.AI cs.CV cs.MM 67%

Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?

Ayan Sar, Sampurna Roy, Tanupriya Choudhury, Ajith Abraham

机构 * School of Computer Sciences, University of Petroleum and Energy Studies (UPES), Dehradun(计算机科学学院,石油与能源研究大学(UPES),德里加恩)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in Foundations and Trends in Signal Processing (#1 in Signal Processing, #3 in Computer Science)

Journal ref Foundations and Trends in Signal Processing (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18179 2025-09-24 cs.CV cs.AI 62%

The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes

Sai Varun Kodathala, Rakesh Vunnam

机构 * Sports Vision, Inc.(体育视觉公司) Vizworld, Inc.(Vizworld公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04545 2025-09-24 cs.CV 57%

PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting

Linqing Wang, Ximing Xing, Yiji Cheng, Zhiyuan Zhao, Donghao Li, Tiankai Hang, Jiale Tao, Qixun Wang, Ruihuang Li, Comi Chen, Xin Li, Mingrui Wu, Xinchi Deng, Shuyang Gu, Chunyu Wang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯文言)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://hunyuan-promptenhancer.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 7 篇

2410.23262 2025-09-24 cs.CV cs.AI cs.CL cs.LG cs.RO 85%

EMMA: End-to-End Multimodal Model for Autonomous Driving

Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, Yin Zhou, James Guo, Dragomir Anguelov, Mingxing Tan

机构 * Waymo LLC(Waymo公司)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by TMLR. Blog post: https://waymo.com/blog/2024/10/introducing-emma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19274 2025-09-24 cs.CL cs.MM 81%

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha

机构 * Indian Institute of Technology Patna(印度理工学院帕纳巴分校) Banasthali Vidyapeeth University(班纳萨利大学) Pandit Deendayal Energy University(德英德能源大学) Manipal University Jaipur(马哈拉施特拉邦大学贾伊普尔分校) Dwarkadas J. Sanghvi College of Engineering(德瓦尔卡斯J.桑格维工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments EMNLP MAINS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18897 2025-09-24 cs.CV 57%

RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing

Jiayu Wang, Ruizhi Wang, Jie Song, Haofei Zhang, Mingli Song, Zunlei Feng, Li Sun

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Software College of Zhejiang University(浙江大学软件学院) Hangzhou City University(杭州市大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 26 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18154 2025-09-24 cs.LG cs.CV 57%

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Tianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang, Wenshuo Ma, Zhihui He, Tianchi Cai, Weize Chen, Yuxiang Huang, Yuanqian Zhao, Bokai Xu, Junbo Cui, Yingjing Xu, Liqing Ruan, Luoyuan Zhang, Hanyu Liu, Jingkun Tang, Hongyuan Liu, Qining Guo, Wenhao Hu, Bingxiang He, Jie Zhou, Jie Cai, Ji Qi, Zonghao Guo, Chi Chen, Guoyang Zeng, Yuxuan Li, Ganqu Cui, Ning Ding, Xu Han, Yuan Yao, Zhiyuan Liu, Maosong Sun

机构 * MiniCPM-V Team, OpenBMB(MiniCPM-V团队,OpenBMB)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Project Website: https://github.com/OpenBMB/MiniCPM-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11101 2025-09-24 cs.CL 57%

Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models

Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li, Peng Zhang

机构 * Tianjin University(天津大学) Shandong Institute of Petroleum and Chemical Technology(山东石油化学工业技术研究所) Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments I need to modify the content of the article

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18599 2025-09-24 q-bio.NC q-bio.QM 50%

From Noise to Insight: Visualizing Neural Dynamics with Segmented SNR Topographies for Improved EEG-BCI Performance

Eva Guttmann-Flury, Shan Zhao, Jian Zhao, Mohamad Sawan

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14139 2025-09-24 cs.CR 50%

Cybersecurity AI: Humanoid Robots as Attack Vectors

Víctor Mayoral-Vilches, Andreas Makris, Kevin Finisterre

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态Agent 6 篇

2509.18405 2025-09-24 cs.CV cs.AI 84%

Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models

Sourav Halder, Jinjun Tong, Xinyu Wu

机构 * U.S. Bank(美国银行)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19087 2025-09-24 cs.CV 79%

Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications

Ganesh Mallya, Yotam Gigi, Dahun Kim, Maxim Neumann, Genady Beryozkin, Tomer Shekel, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19168 2025-09-24 cs.RO 78%

A Multimodal Stochastic Planning Approach for Navigation and Multi-Robot Coordination

Mark Gonzales, Ethan Oh, Joseph Moore

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 Pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18576 2025-09-24 cs.RO cs.AI 77%

LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA

Zeyi Kang, Liang He, Yanxin Zhang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学中国) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学巴黎法国)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18372 2025-09-24 cs.CV 57%

TinyBEV: Cross Modal Knowledge Distillation for Efficient Multi Task Bird's Eye View Perception and Planning

Reeshad Khan, John Gauch

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18609 2025-09-24 cs.RO 50%

PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving

Chengran Yuan, Zijian Lu, Zhanqi Zhang, Yimin Zhao, Zefan Huang, Shuo Sun, Jiawei Sun, Jiahui Li, Christina Dao Wen Lee, Dongen Li, Marcelo H. Ang

机构 * Department of Mechanical Engineering, National University of Singapore(机械工程系,国立新加坡大学) Department of Civil and Environmental Engineering, National University of Singapore(土木与环境工程系,国立新加坡大学)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态训练与对齐 14 篇

2509.19212 2025-09-24 cs.CL cs.AI 84%

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang

机构 * University of Notre Dame(notre dame 大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments A lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12623 2025-09-24 cs.SD cs.AI cs.CL cs.MM eess.AS 83%

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 83%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 82%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17492 2025-09-24 cs.CV cs.AI 81%

Multimodal Medical Image Classification via Synergistic Learning Pre-training

Qinghua Lin, Guang-Hai Liu, Zuoyong Li, Yang Li, Yuting Jiang, Xiang Wu

机构 * College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院) College of Computer Science and Engineering, Guangxi Normal University(广西师范大学计算机科学与工程学院) Fujian Provincial Key Laboratory of Information Processing and Intelligent Control, School of Computer and Big Data, Minjiang University(福建省信息处理与智能控制重点实验室,闽江学院计算机与大数据学院) Department of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) Department of Digestive Endoscopy, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院消化内镜科,福建医科大学省临床医学学院) Department of Urology, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院泌尿科,福建医科大学省临床医学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19208 2025-09-24 cs.CV 79%

Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data

Earl Ranario, Ismael Mayanja, Heesup Yun, Brian N. Bailey, J. Mason Earles

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19047 2025-09-24 cs.RO 67%

ManipForce: Force-Guided Policy Learning with Frequency-Aware Representation for Contact-Rich Manipulation

Geonhyup Lee, Yeongjin Lee, Kangmin Kim, Seongju Lee, Sangjun Noh, Seunghyeok Back, Kyoobin Lee

机构 * Department of AI Convergence, Gwangju Institute of Science and Technology (GIST)(人工智能融合系,全州科学技术院) Department of AI Machinery, Korea Institute of Machinery & Materials (KIMM)(人工智能机械系,韩国机械材料研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04183 2025-09-24 cs.CL cs.AI 62%

GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding

Ziyin Zhang, Hang Yu, Shijie Li, Peng Di, Jianguo Li, Rui Wang

机构 * Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18743 2025-09-24 cs.CV 57%

TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing

Susmit Neogi

机构 * Department of Mechanical Engineering(机械工程系) Indian Institute of Technology Bombay(印度理工学院班加罗尔) Mumbai, India(孟买,印度)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18738 2025-09-24 cs.CV 57%

HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection

Ruichao Hou, Xingyuan Li, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18733 2025-09-24 cs.CV 57%

Knowledge Transfer from Interaction Learning

Yilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu, Shugong Xu

机构 * Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏