arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-01 至 2025-08-01 共收录 46 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2411.18659 2025-08-01 cs.CV cs.AI 84%

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学、腾讯) Tencent(腾讯) Tencent, University of Macau(腾讯、澳门大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15621 2025-08-01 cs.CV cs.AI cs.CL cs.MM 74%

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) IIT-CNR(意大利国家研究 council(IIT))

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments ICCV 2025 Workshop on What is Next in Multimodal Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19480 2025-08-01 cs.CV 57%

GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers

Shijie Ma, Yuying Ge, Teng Wang, Yuxin Guo, Yixiao Ge, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Institute of Automation, CAS(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025. Project released at: https://mashijie1028.github.io/GenHancer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23362 2025-08-01 cs.CV 57%

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers

Ji Ma, Wei Suo, Peng Wang, Yanning Zhang

机构 * Northwestern Polytechnical University(西北工业大学) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted By ACM MM 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23226 2025-08-01 cs.CV 57%

Toward Safe, Trustworthy and Realistic Augmented Reality User Experience

Yanming Xiu

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程系,杜克大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 2 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2507.22886 2025-08-01 cs.CV 85%

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Kaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang Jiang

机构 * Fudan University, China(复旦大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025, Project Page: https://henghuiding.com/OmniAVS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23010 2025-08-01 cs.LG cs.AI cs.CV cs.SD eess.AS 82%

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods

Siwoo Park

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22934 2025-08-01 cs.CL cs.AI 81%

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

Jingwei Zhao, Yuhua Wen, Qifei Li, Minchi Hu, Yingying Zhou, Jingyao Xue, Junyang Wu, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23544 2025-08-01 cs.RO cs.CV cs.HC 79%

User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals

Ryo Miyoshi, Yuki Okafuji, Takuya Iwamoto, Junya Nakanishi, Jun Baba

机构 * CyberAgent Osaka University(大阪大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted for presentation at IEEE/RSJ International Conference on Intelligent Robots and Systems 2025 (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08052 2025-08-01 eess.AS 79%

Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

Ming Cheng, Ming Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by IEEE Transactions on Audio, Speech, and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23590 2025-08-01 cs.SD eess.AS 57%

Identifying Hearing Difficulty Moments in Conversational Audio

Jack Collins, Adrian Buzea, Chris Collier, Alejandro Ballesta Rosen, Julian Maclaren, Richard F. Lyon, Simon Carlile

机构 * Google Research Australia(谷歌澳大利亚研究)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 2 篇

2304.01430 2025-08-01 cs.CV cs.AI cs.LG 62%

Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots

Dong Lao, Zhengyang Hu, Francesco Locatello, Yanchao Yang, Stefano Soatto

机构 * UCLA(加州大学洛杉矶分校) HKU(香港大学) ISTA(因斯布鲁克大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20947 2025-08-01 cs.CV cs.MM 62%

Hierarchical Sub-action Tree for Continuous Sign Language Recognition

Dejie Yang, Zhu Xu, Xinjie Gao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Peking Universitys, Beijing, China(通用人工智能国家重点实验室,北京大学,北京,中国)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2507.22896 2025-08-01 cs.HC cs.AI cs.CV cs.RO 84%

iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement

Kohou Wang, ZhaoXiang Liu, Lin Bai, Kun Fan, Xiang Liu, Huan Hu, Kai Wang, Shiguo Lian

机构 * Unicom Data Intelligence(中国联通数据智能研究所) Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院) China United Network Communications Group Corporation Limited(中国联合网络通信集团有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23331 2025-08-01 cs.CV 83%

Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision

Qiang Lu, Waikit Xiu, Xiying Li, Shenyu Hu, Shengbo Sun

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) Guangdong Provincial Key Laboratory of Intelligent Transportation System(广东省智能交通系统重点实验室)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23188 2025-08-01 cs.CV 79%

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian, Mingyuan Zhang, Weixin Si, Shihao Zou

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海科技大学) Nanyang Technological University(南洋理工大学) Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology(计算机科学与控制工程学院,深圳大学先进技术学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22938 2025-08-01 cs.CL cs.AI 76%

A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents

Sumit Soman, H. G. Ranjani, Sujoy Roychowdhury, Venkata Dharma Surya Narayana Sastry, Akshat Jain, Pranav Gangrade, Ayaaz Khan

机构 * Ericsson R&D Bangalore Karnataka India(爱立信研发部班加罗尔卡纳塔克邦印度)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Comments Accepted for publication at the KDD 2025 Workshop on Structured Knowledge for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23217 2025-08-01 cs.LG cs.AI 57%

Zero-Shot Document Understanding using Pseudo Table of Contents-Guided Retrieval-Augmented Generation

Hyeon Seong Jeong, Sangwoo Jo, Byeong Hyun Yoon, Yoonseok Heo, Haedong Jeong, Taehoon Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2507.23058 2025-08-01 cs.CV cs.AI 81%

Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

Alexandru Buburuzan

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments A dissertation submitted to The University of Manchester for the degree of Bachelor of Science in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22920 2025-08-01 cs.CL cs.AI 81%

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey

Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23202 2025-08-01 cs.CV 79%

Adversarial-Guided Diffusion for Multimodal LLM Attacks

Chengwei Xia, Fan Ma, Ruijie Quan, Kun Zhan, Yi Yang

机构 * School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21872 2025-08-01 cs.AI 79%

MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

Shouyi Lu, Zihan Lin, Chao Lu, Huanran Wang, Guirong Zhuo, Lianqing Zheng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23676 2025-08-01 cs.LG cs.CV 74%

DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data

Rabeya Tus Sadia, Qiang Cheng

机构 * Department of Computer Science University of Kentucky(计算机科学系 哥伦比亚大学) Department of Computer Science, Institute for Biomedical Informatics University of Kentucky(计算机科学系 生物医学信息学研究所 哥伦比亚大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23178 2025-08-01 cs.SE cs.AI 57%

AutoBridge: Automating Smart Device Integration with Centralized Platform

Siyuan Liu, Zhice Yang, Huangxun Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ShanghaiTech University(上海科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 14 pages, 12 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22789 2025-08-01 cs.LG cs.AI 57%

G-Core: A Simple, Scalable and Balanced RLHF Trainer

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Haoqiang Hong, Boqi Liu, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments I haven't received company approval yet, and I uploaded it by mistake

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17761 2025-08-01 cs.CV 57%

Step1X-Edit: A Practical Framework for General Image Editing

Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, Guopeng Li, Yuang Peng, Quan Sun, Jingwei Wu, Yan Cai, Zheng Ge, Ranchen Ming, Lei Xia, Xianfang Zeng, Yibo Zhu, Binxing Jiao, Xiangyu Zhang, Gang Yu, Daxin Jiang

机构 * Step1X-Image Team(Step1X-图像团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments code: https://github.com/stepfun-ai/Step1X-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05422 2025-08-01 cs.CV cs.LG cs.RO 57%

EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction via Polynomial Representations

Yue Yao, Mohamed-Khalil Bouzidi, Daniel Goehring, Joerg Reichardt

机构 * Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学) Continental Automotive GmbH(大陆汽车股份有限公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2507.23382 2025-08-01 cs.CL cs.AI cs.CV 85%

MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

Yiyan Ji, Haoran Chen, Qiguang Chen, Chengyue Wu, Libo Qin, Wanxiang Che

机构 * Harbin Institute of Technology(哈尔滨工业大学) The University of Hong Kong(香港大学) Central South University(中南大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10006 2025-08-01 cs.MM cs.AI cs.CV cs.LG 82%

HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction

Jie Qin, Wei Yang, Yan Su, Yiran Zhu, Weizhen Li, Yunyue Pan, Chengchang Pan, Honggang Qi

机构 * School of Computer Science and Technology, University of the Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒实验室)

专题命中 多模态评测 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 8 pages,6 figures,3 tables,accepted by the 33rd ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23135 2025-08-01 cs.CL 79%

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Ananya Sadana, Yash Kumar Lal, Jiawei Zhou

机构 * Stony Brook University(石溪大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏