arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-15 至 2025-08-15 共收录 59 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 10 篇

2508.01272 2025-08-15 cs.CV 57%

PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation

Zonglei Jing, Xiao Yang, Xiaoqian Li, Siyuan Liang, Aishan Liu, Mingchuan Zhang, Xianglong Liu

机构 * Beihang University(北航) Beijing University of Posts and Telecommunications(北京邮电大学) Taishan University(泰山大学) Nanyang Technological University(南洋理工大学) Henan University of Science and Technology(河南科技大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19948 2025-08-15 cs.RO 50%

Motion Planning Diffusion: Learning and Adapting Robot Motion Planning with Diffusion Models

J. Carvalho, A. Le, P. Kicki, D. Koert, J. Peters

机构 * Intelligent Autonomous Systems Lab, Computer Science Department, Technical University of Darmstadt, Germany(德意志图林根技术大学计算机科学系智能自主系统实验室) Poznan University of Technology, Poland(波兰波兹南技术大学) IDEAS, Warsaw, Poland(波兰华沙IDEAS) Centre for Cognitive Science, Technical University of Darmstadt, Germany(德意志图林根技术大学认知科学中心) German Research Center for AI (DFKI), Research Department: SAIROL, Darmstadt, Germany(德国人工智能研究中心(DFKI)研究部:SAIROL,德意志图林根,德国) Hessian.AI, Darmstadt, Germany(黑森AI,德意志图林根,德国)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 15 篇

2508.06800 2025-08-15 cs.LG cs.AI 83%

Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing Modalities

Rui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan, Qi Fan

机构 * Inner Mongolia University(内蒙古大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10015 2025-08-15 cs.CL 83%

RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis

Enzhi Wang, Qicheng Li, Shiwan Zhao, Aobo Kong, Jiaming Zhou, Xi Yang, Yequan Wang, Yonghua Lin, Yong Qin

机构 * TMCC, College of Computer Science, Nankai University(TMCC,计算机科学学院,南开大学) Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09999 2025-08-15 cs.CL cs.LG 83%

XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

Yuzhuo Xiao, Zeyu Han, Yuhan Wang, Huaizu Jiang

机构 * Guizhou University(贵州大学) Northeastern University(东北大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments For associated code and dataset, see https://github.com/neu-vi/XFacta

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06869 2025-08-15 cs.CV cs.LG 83%

CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language Models

Junho Kim, Hyungjin Chung, Byung-Hoon Kim

机构 * EverEx Yonsei University(延世大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10769 2025-08-15 cs.AI cs.MM 81%

Modeling Human Responses to Multimodal AI Content

Zhiqi Shen, Shaojing Fan, Danni Xu, Terence Sim, Mohan Kankanhalli

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10655 2025-08-15 cs.CV cs.AI 81%

Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and Benchmarking

Zhangyong Tang, Tianyang Xu, Xuefeng Zhu, Chunyang Cheng, Tao Zhou, Xiaojun Wu, Josef Kittler

机构 * Jiangnan University(江南大学) Nanjing University of Science and Technology(南京理工大学) University of Surrey(萨里大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10429 2025-08-15 cs.AI cs.CR cs.CV 81%

MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance

Yi Dong, Yusuke Muraoka, Scott Shi, Yi Zhang

机构 * Codatta Kite AI Binance Wallet Community Codatta Community

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures, 6 tables. The dataset is available at https://huggingface.co/datasets/Codatta/MM-Food-100K

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04801 2025-08-15 cs.CV 79%

OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM

Jinhong Wang, Shuo Tong, Jian liu, Dongqi Tang, Weiqiang Wang, Wentong Li, Hongxia Xu, Danny Chen, Jintai Chen, Jian Wu

机构 * College of Computer Science & Technology, Zhejiang University(浙江大学计算机科学与技术学院) Transvascular Implantation Devices Research Institute and Liangzhu Laboratory(血管植入物研究机构和良渚实验室) Ant Group(蚂蚁集团) University of Notre Dame(圣母大学) HKUST (Guangzhou)(香港科技大学(广州))

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10468 2025-08-15 cs.HC 78%

Stress Detection from Multimodal Wearable Sensor Data

Paul Schreiber, Beyza Cinar, Lennart Mackert, Maria Maleshkova

专题命中 多模态评测 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03497 2025-08-15 cs.CV 74%

EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware Evaluation

Deqiang Yin, Junyi Guo, Huanda Lu, Fangyu Wu, Dongming Lu

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Zhejiang University(浙江大学)

专题命中 多模态评测 :MLLM(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10869 2025-08-15 cs.CV cs.AI 62%

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Sushant Gautam, Vajira Thambawita, Michael Riegler, Pål Halvorsen, Steven Hicks

机构 * SimulaMet - Simula Metropolitan Center for Digital Engineering, Oslo, Norway(SimulaMet - Simula Metropolitan Center for Digital Engineering,挪威奥斯陆) Simula Research Laboratory, Oslo, Norway(Simula研究实验室,挪威奥斯陆) OsloMet - Oslo Metropolitan University, Oslo, Norway(OsloMet - 奥斯陆 Metropolitan 大学,挪威奥斯陆)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10771 2025-08-15 cs.CV cs.AI 62%

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences

Jieyu Li, Xin Zhang, Joey Tianyi Zhou

机构 * National University of Singapore(新加坡国立大学) Agency for Science, Technology and Research(科技研究局)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Proceedings of the 33rd ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10433 2025-08-15 cs.AI cs.CV cs.LG 62%

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Runqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang, Xiaowan Wang, Enhui Wan, Sitong Zhou, Guanting Dong, Yuchen Zeng, Yida Xu, Jie Wang, Chong Sun, Chen Li, Honggang Zhang

机构 * BUPT(北京邮电大学) WeChat Vision, Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10806 2025-08-15 cs.AI 57%

Who Benefits from AI Explanations? Towards Accessible and Interpretable Systems

Maria J. P. Peixoto, Akriti Pandey, Ahsan Zaman, Peter R. Lewis

机构 * Ontario Tech University(安大略技术大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Paper accepted for the IJCAI 2025 Workshop on Explainable Artificial Intelligence (XAI): https://sites.google.com/view/xai2025/proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10770 2025-08-15 cs.CV 57%

From Diagnosis to Improvement: Probing Spatio-Physical Reasoning in Vision Language Models

Tiancheng Han, Yunfei Gao, Yong Li, Wuzhou Yu, Qiaosheng Zhang, Wenqi Shao

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 1 篇

2507.10778 2025-08-15 cs.CV cs.AI 73%

Warehouse Spatial Question Answering with LLM Agent

Hsiang-Wei Huang, Jen-Hao Cheng, Kuang-Ming Chen, Cheng-Yen Yang, Bahaa Alattar, Yi-Ru Lin, Pyongkun Kim, Sangwon Kim, Kwangju Kim, Chung-I Huang, Jenq-Neng Hwang

机构 * Information Processing Lab, University of Washington, USA(华盛顿大学信息处理实验室) Electronics and Telecommunications Research Institute, South Korea(韩国电子电信研究院) National Center for High-performance Computing, Taiwan(台湾高性能计算国家研究中心)

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 1st Place Solution of the 9th AI City Challenge Track 3

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 9 篇

2508.10371 2025-08-15 cs.RO 85%

Few-shot Vision-based Human Activity Recognition with MLLM-based Visual Reinforcement Learning

Wenqi Zheng, Yutaka Arakawa

机构 * Graduate School(研究生院) Faculty of Information Science(信息科学系) Electrical Engineering(电气工程) Kyushu University(九州大学) JAPAN(日本)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06293 2025-08-15 cs.CV 83%

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness

Qifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu, Bosheng Qin, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10552 2025-08-15 cs.CL cs.AI 81%

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Huyu Wu, Meng Tang, Xinhan Zheng, Haiyun Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10753 2025-08-15 cs.IR 78%

Hypercomplex Prompt-aware Multimodal Recommendation

Zheyu Chen, Jinfeng Xu, Hewei Wang, Shuo Yang, Zitong Wan, Haibo Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10696 2025-08-15 cs.CE 78%

Chem3DLLM: 3D Multimodal Large Language Models for Chemistry

Lei Jiang, Shuzhou Sun, Biqing Qi, Yuchen Fu, Xiaohua Xu, Yuqiang Li, Dongzhan Zhou, Tianfan Fu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10116 2025-08-15 cs.IR 67%

Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment

Yipeng Zhang, Hongju Yu, Aritra Mandal, Canran Xu, Qunzhi Zhou, Zhe Wu

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10704 2025-08-15 cs.CV 57%

Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios

Zhanwen Liu, Yujing Sun, Yang Wang, Nan Yang, Shengbo Eben Li, Xiangmo Zhao

机构 * School of Information Engineering, Chang’an University(信息工程学院,长安大学) School of Vehicle Mobility & College of AI, Tsinghua University(车辆运动学院与人工智能学院,清华大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10232 2025-08-15 cs.CV 57%

CellSymphony: Deciphering the molecular and phenotypic orchestration of cells with single-cell pathomics

Paul H. Acosta, Pingjun Chen, Simon P. Castillo, Maria Esther Salvatierra, Yinyin Yuan, Xiaoxi Pan

机构 * Translational Molecular Pathology Department, Division of Pathology and Laboratory Medicine, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心转化分子病理学部门) Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心肿瘤数据科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11049 2025-08-15 cs.LG cs.AI 57%

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning

Andrew P. Berg, Qian Zhang, Mia Y. Wang

机构 * Department of Computer Science(计算机科学系) College of Charleston(查尔斯顿学院) Department of Engineering(工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他多模态 2 篇

2507.09697 2025-08-15 physics.optics eess.IV 50%

Curvature-adaptive gigapixel microscopy at submicron resolution and centimeter scale

Xi Yang, Haitao Chen, Lucas Kreiss, Clare B. Cook, Genevieve Kuczewski, Mark Harfouche, Martin O. Bohlen, Roarke Horstmeyer

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15604 2025-08-15 astro-ph.GA 50%

Evidence for Multiple Types of Post-Starburst Galaxies

Emma W. Nielsen, Charles L. Steinhardt, Mathieux Harper, Conor McPartland, Aidan Sedgewick

专题命中 其他多模态 :multimodal(abstract)

Comments 8 pages, 8 figures

Journal ref A&A 700, A116 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏