arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-30 至 2025-07-30 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2506.09081 2025-07-30 cs.CV cs.AI cs.CL 85%

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li, Jin-Ge Yao, Bowen Qin, Richeng Xuan, Xi Yang

机构 * BAAI FlagEval Team(BAAI 评测团队)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20217 2025-07-30 cs.RO cs.AI cs.CV 84%

Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

Wei Cui, Haoyu Wang, Wenkang Qin, Yijie Guo, Gang Han, Wen Zhao, Jiahang Cao, Zhang Zhang, Jiaru Zhong, Jingkai Sun, Pihai Sun, Shuai Shi, Botuo Jiang, Jiahao Ma, Jiaxu Wang, Hao Cheng, Zhichao Liu, Yang Wang, Zheng Zhu, Guan Huang, Jian Tang, Qiang Zhang

机构 * X-Humanoid GigaAI Project(GigaAI项目)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21924 2025-07-30 cs.CV 79%

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning

Tianhong Gao, Yannian Fu, Weiqun Wu, Haixiao Yue, Shanshan Liu, Gang Zhang

机构 * Baidu Inc.(百度公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16876 2025-07-30 q-bio.QM cs.AI cs.LG 79%

Machine learning-based multimodal prognostic models integrating pathology images and high-throughput omic data for overall survival prediction in cancer: a systematic review

Charlotte Jennings, Andrew Broad, Lucy Godson, Emily Clarke, David Westhead, Darren Treanor

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Main article (50 pages, inc 3 tables, 4 figures). Supplementary material included with additional methodological information and data

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 70%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21157 2025-07-30 cs.CR cs.CV 70%

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

Naseem Khan, Tuan Nguyen, Amine Bermak, Issa Khalil

机构 * Department of Computer Science(计算机科学系) Hamad bin Khalifa University(哈马德·本·哈利法大学) Qatar Computing Research Institute(卡塔尔计算研究所)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 27 pages, 4 Tables, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21104 2025-07-30 cs.CL cs.AI 62%

iLSU-T: an Open Dataset for Uruguayan Sign Language Translation

Ariel E. Stassi, Yanina Boria, J. Matías Di Martino, Gregory Randall

机构 * Universidad de la República(乌拉圭共和国大学) Universidad de Buenos Aires(布宜诺斯艾利斯大学) Universidad Católica del Uruguay(乌拉圭天主教大学) Duke University(杜克大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, 19th International Conference on Automatic Face and Gesture Recognition IEEE FG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21335 2025-07-30 cs.CV 57%

Analyzing the Sensitivity of Vision Language Models in Visual Question Answering

Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal

机构 * University of Memphis, TN, USA(密苏里大学) Adobe Research, San Jose, CA, USA(Adobe研究院) University of Maryland Baltimore County, USA(马里兰大学巴尔的摩县分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21520 2025-07-30 cs.IR 50%

Solution for Meta KDD Cup'25: A Comprehensive Three-Step Framework for Vision Question Answering

Zijian Zhang, Xiaocheng Zhang, Yang Zhou, Zhimin Lin, Peng Yan

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2507.21199 2025-07-30 cs.LG cs.AI cs.DC cs.HC 79%

Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications

Xinye Cao, Hongcan Guo, Guoshun Nan, Jiaoyang Cui, Haoting Qian, Yihan Lin, Yilin Peng, Diyang Zhang, Yanzhao Hou, Huici Wu, Xiaofeng Tao, Tony Q. S. Quek

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted by IEEE JSAC. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18084 2025-07-30 cs.RO cs.AI 79%

GSON: A Group-based Social Navigation Framework with Large Multimodal Model

Shangyi Luo, Peng Sun, Ji Zhu, Yuhong Deng, Cunjun Yu, Anxing Xiao, Xueqian Wang

机构 * Center for Artificial Intelligence and Robotics, Tsinghua Shenzhen International Graduate School(人工智能与机器人中心,清华大学深圳国际研究生院) School of Computing, National University of Singapore(computing 学院,新加坡国立大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted by IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21454 2025-07-30 eess.SP 50%

Transmission With Machine Language Tokens: A Paradigm for Task-Oriented Agent Communication

Zhuoran Xiao, Chenhui Ye, Yijia Feng, Yunbo Hu, Tianyu Jiao, Liyu Cai, Guangyi Liu

专题命中 多模态Agent :multi-modal(abstract)

Comments Accepted by IEEE Globecom 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 7 篇

2507.21395 2025-07-30 cs.MM cs.AI cs.SD eess.AS 89%

Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

Zeyu Deng, Yanhui Lu, Jiashu Liao, Shuang Wu, Chongfeng Wei

机构 * James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) University of Bristol(布里斯托大学) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院) School of Computing Science, University of Glasgow(格拉斯哥大学计算科学学院) Department of Civil, Environmental & Geomatic Engineering, University College London (UCL)(伦敦大学学院(UCL)土木、环境与测绘工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08204 2025-07-30 cs.CV 83%

One-stage Modality Distillation for Incomplete Multimodal Learning

Shicai Wei, Yang Luo, Chunbo Luo

机构 * School of Information and Communication Engineering University of Electronic Science and Technology of China(信息与通信工程学院 电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03248 2025-07-30 cs.CV cs.AI cs.CL 82%

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Yiwu Zhong, Zhuoming Liu, Yin Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 79%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10774 2025-07-30 cs.LG cs.AI 79%

Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting

Yueyang Yao, Jiajun Li, Xingyuan Dai, MengMeng Zhang, Xiaoyan Gong, Fei-Yue Wang, Yisheng Lv

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10091 2025-07-30 eess.IV 78%

G$^{2}$SF-MIAD: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly Detection

Chengyu Tao, Xuanming Cao, Juan Du

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21857 2025-07-30 cs.CV 57%

Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection

Jiahao He, Daerji Suolang, Keren Fu, Qijun Zhao

机构 * College of Computer Science, Sichuan University(四川大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments submitted to TMM on 11-Jun-2024, ID: MM-020522, still in peer review

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 3 篇

2507.21378 2025-07-30 cs.HC cs.AI 79%

ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices

Kevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag, Raj Sodhi, Tanya Jonker

机构 * University of Toronto(多伦多大学) Meta Reality Labs(Meta现实实验室)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to UIST'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.06871 2025-07-30 stat.CO stat.ML 78%

Sampling from high-dimensional, multimodal distributions using automatically tuned, tempered Hamiltonian Monte Carlo

Joonha Park

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21158 2025-07-30 cs.AI cs.HC 74%

Adaptive XAI in High Stakes Environments: Modeling Swift Trust with Multimodal Feedback in Human AI Teams

Nishani Fernando, Bahareh Nakisa, Adnan Ahmad, Mohammad Naim Rastgoo

机构 * Deakin University(德金大学) Monash University(莫纳什大学)

专题命中 其他多模态 :multimodal(title);分类 cs.AI

Comments 15 pages, 1 figure, Accepted to MAI-XAI@ECAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏