arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-09 至 2025-09-09 共收录 75 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2509.05925 2025-09-09 cs.CV cs.IT math.IT 88%

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

Ruiqi Shen, Haotian Wu, Wenjing Zhang, Jiangjing Hu, Deniz Gunduz

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院电子与电气工程系)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Published as a conference paper at IEEE 35th Workshop on Machine Learning for Signal Processing (MLSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08040 2025-09-09 cs.LG cs.AI 79%

BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models

Maozhen Zhang, Mengnan Zhao, Wei Wang, Bo Wang

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学) New Laboratory of Pattern Recognition (NLPR) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)多模态人工智能系统国家重点实验室(MAIS)自动化研究所,中国科学院(CASIA))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05786 2025-09-09 cs.MM cs.SD eess.AS 73%

Effectively obtaining acoustic, visual and textual data from videos

Jorge E. León, Miguel Carrasco

机构 * Adolfo Ibánez University(阿多洛·伊巴涅斯大学) Diego Portales University(迪埃戈·波特莱斯大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06535 2025-09-09 cs.CV cs.AI cs.LG 62%

On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''

Hua Chang Bakker, Stan Fris, Angela Madelon Bernardy, Stan Deutekom

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06759 2025-09-09 cs.LG cs.AI 57%

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization

Thanh Thi Nguyen, Campbell Wilson, Janis Dalins

机构 * AiLECS Lab, Monash University Melbourne, Australia(墨尔本大学AiLECS实验室,澳大利亚) AiLECS Lab, Monash University, Australia ICMEC Australia, Sydney, Australia(墨尔本大学AiLECS实验室,澳大利亚 ICMEC澳大利亚,悉尼,澳大利亚)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05669 2025-09-09 cs.CV 57%

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance

Weijie Shen, Xinrui Wang, Yuanqi Nie, Apiradee Boonmee

机构 * Beihua University(白华大学) Kasem Bundit University(Kasem Bundit大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19757 2025-09-09 cs.RO cs.CV 57%

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Zhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao, Jifeng Dai, Yuntao Chen

机构 * Shanghai AI Lab(上海人工智能实验室) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) Peking University(北京大学) SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) HKISI, CAS(中国科学院香港中文大学研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Preprint; https://robodita.github.io; To appear in ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10032 2025-09-09 cs.CV 57%

Osprey: Pixel Understanding with Visual Instruction Tuning

Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, Jianke Zhu

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Microsoft(微软) The HongKong Polytechnical University(香港理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR2024, Code and Demo link:https://github.com/CircleRadon/Osprey

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06768 2025-09-09 cs.RO 50%

Embodied Hazard Mitigation using Vision-Language Models for Autonomous Mobile Robots

Oluwadamilola Sotomi, Devika Kodi, Kiruthiga Chandra Shekar, Aliasghar Arab

机构 * Department of Mechanical and Aerospace Engineering, Tandon School of Engineering, New York University(机械与航空航天工程系,坦顿工程学院,纽约大学) GenAuto.ai by General Autonomy Inc.(General Autonomy Inc. 的 GenAuto.ai)

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2402.12226 2025-09-09 cs.CL cs.AI cs.CV cs.LG 85%

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu

机构 * Fudan University(复旦大学) Multimodal Art Projection Research Community(多模态艺术投影研究社区) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 16 figures, under review, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06074 2025-09-09 cs.CL 79%

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li

机构 * Inner Mongolia University(内蒙古大学) Center for Language and Speech Processing (CLSP)(语言与语音处理中心) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06598 2025-09-09 eess.AS cs.AI cs.LG eess.IV eess.SP 79%

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI、eess.AS

Comments arXiv admin note: substantial text overlap with arXiv:2507.04845

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06382 2025-09-09 cs.HC 78%

Context-Adaptive Hearing Aid Fitting Advisor through Multi-turn Multimodal LLM Conversation

Yingke Ding, Zeyu Wang, Xiyuxing Zhang, Hongbin Chen, Zhenan Xu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Ubicomp Companion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 11 篇

2411.19787 2025-09-09 cs.LG cs.AI 83%

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

Armin Saghafian, Amirmohammad Izadi, Negin Hashemi Dijujin, Mahdieh Soleymani Baghshah

机构 * Sharif University of Technology(谢里夫理工大学)

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06831 2025-09-09 cs.CV 79%

Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis

Simon Pezold, Jérôme A. Kurylec, Jan S. Liechti, Beat P. Müller, Joël L. Lavanchy

机构 * Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(巴塞尔大学生物医学工程系) Clarunis – University Digestive Health Care Center Basel, Basel, Switzerland(Clarunis – 巴塞尔大学消化健康医疗中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 3 figures; accepted at ML-CDS @ MICCAI 2025, Daejeon, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06422 2025-09-09 cs.CV 79%

Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM

Hua Zhang, Changjiang Luo, Ruoyu Chen

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06389 2025-09-09 cs.SD cs.AI 79%

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Xiaoran Yang, Jianxuan Yang, Xinyue Guo, Haoyu Wang, Ningning Pan, Gongping Huang

机构 * School of Electronic Information, Wuhan University, Wuhan, China(武汉大学电子信息学院) MiLM Plus, Xiaomi Inc., Wuhan, China(小米公司MiLM Plus团队)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05898 2025-09-09 cs.HC 78%

Attention, Action, and Memory: How Multi-modal Interfaces and Cognitive Load Alter Information Retention

Omar Elgohary, Zhu-Tien

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19153 2025-09-09 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 73%

QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning

Yinuo Wang, Gavin Tao

机构 * Allen Wang Gavin Tao

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14pages, 9 figures, Journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09632 2025-09-09 cs.CV cs.AI 62%

Preacher: Paper-to-Video Agentic System

Jingwei Liu, Ling Yang, Hao Luo, Fan Wang, Hongyan Li, Mengdi Wang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) Hupan Lab(虎扑实验室) National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05298 2025-09-09 cs.HC cs.AI cs.MM 62%

Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression

Rui Xi, Xianghan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted to the Proceedings of the 2025 International Conference on Artificial Intelligence and Virtual Reality (AIVR 2025). \c{opyright} 2025 Springer. This is the author-accepted manuscript. Rui Xi and Xianghan Wang contributed equally to this work. The final version will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06023 2025-09-09 cs.CV 57%

DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion

Mengmeng Liu, Michael Ying Yang, Jiuming Liu, Yunpeng Zhang, Jiangtao Li, Sander Oude Elberink, George Vosselman, Hao Cheng

机构 * University of Twente(特文特大学) University of Bath(巴斯大学) Shanghai Jiao Tong University(上海交通大学) PhiGent Robotics(PhiGent机器人)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01563 2025-09-09 cs.CV 57%

Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Guowang Zhang, Han Shen, Hao Peng, Haojie Ding, Hao Wang, Haonan Fan, Hengrui Ju, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Muhao Wei, Qiang Wang, Ruitao Wang, Sen Na, Shengnan Zhang, Siyang Mao, Sui Huang, Tianke Zhang, Tingting Gao, Wei Chen, Wei Yuan, Xiangyu Wu, Xiao Hu, Xingyu Lu, Yi-Fan Zhang, Yiping Yang, Yulong Chen, Zeyi Lu, Zhenhua Wu, Zhixin Ling, Zhuoran Yang, Ziming Li, Di Xu, Haixuan Gao, Hang Li, Jing Wang, Lejian Ren, Qigen Hu, Qianqian Wang, Shiyao Wang, Xinchen Luo, Yan Li, Yuhang Hu, Zixing Zhang

机构 * Keye Team, Kuaishou Group(快手集团Keye团队)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Github page: https://github.com/Kwai-Keye/Keye

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09423 2025-09-09 cs.RO cs.CV 57%

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Kechun Xu, Xunlong Xia, Kaixuan Wang, Yifei Yang, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang

机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) Alibaba Cloud(阿里云) Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by T-ASE and CoRL25 GenPriors Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2509.05883 2025-09-09 cs.CR cs.AI 74%

Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs

Andrew Yeo, Daeseon Choi

机构 * Ranchview High School(拉文斯维尔高中) Soongsil University(松山大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments 8 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06566 2025-09-09 cs.CV 57%

Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval

Emil Demić, Luka Čehovin Zajc

机构 * Faculty of Computer and Information Science University of Ljubljana(计算机与信息科学系卢布尔雅那大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to BMVC2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2506.07963 2025-09-09 cs.AI cs.CL cs.CV 82%

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) School of Computer Science(计算机科学学院) Baidu Inc.(百度公司) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05263 2025-09-09 cs.AI cs.CV cs.LG 81%

LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation

Yinglin Duan, Zhengxia Zou, Tongwei Gu, Wei Jia, Zhan Zhao, Luyi Xu, Xinzhu Liu, Yenan Lin, Hao Jiang, Kang Chen, Shuang Qiu

机构 * NetEase, Inc.(网易公司) Beihang University(北京航空航天大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学) Independent Researcher & Technical Artists(独立研究者及技术艺术家)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05714 2025-09-09 cs.AI cs.CV 81%

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs

Zhaoyu Fan, Kaihang Pan, Mingze Zhou, Bosheng Qin, Juncheng Li, Shengyu Zhang, Wenqiao Zhang, Siliang Tang, Fei Wu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12528 2025-09-09 cs.CV 79%

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏