arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-06 至 2025-08-06 共收录 64 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2502.12591 2025-08-06 cs.CV cs.CL 84%

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao, Thong Nguyen, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) National University of Singapore, Singapore(国立新加坡大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12513 2025-08-06 cs.CV 83%

RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm

Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, Jiankang Deng

机构 * The University of Sydney(悉尼大学) DeepGlint Imperial College London(伦敦帝国学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 15 pages, 12 figures, Accepted by ACM MM2025, Webpage: https://garygutc.github.io/RealSyn

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03654 2025-08-06 cs.CL cs.CV 81%

Can Large Vision-Language Models Understand Multimodal Sarcasm?

Xinyu Wang, Yue Zhang, Liqiang Jing

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02890 2025-08-06 cs.CV cs.CL 62%

VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction

Rongxin Jiang, Robert Long, Chenghao Gu, Mingrui Yan

机构 * Heilongjiang University of Science and Technology(黑龙江科技大学) University of Padua(帕多瓦大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2508.02849 2025-08-06 eess.AS cs.AI cs.CL cs.SD 85%

SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec

Chunyu Qiang, Haoyu Wang, Cheng Gong, Tianrui Wang, Ruibo Fu, Tao Wang, Ruilong Chen, Jiangyan Yi, Zhengqi Wen, Chen Zhang, Longbiao Wang, Jianwu Dang, Jianhua Tao

机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) Kuaishou Technology(快手科技) Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University(认知计算与应用重点实验室,智能计算学院,天津大学) Department of Automation, BNRist, Tsinghua University(自动化系,北研所,清华大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Science(深圳先进技术研究院,中国科学院)

专题命中 音频语音多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02905 2025-08-06 cs.CV cs.SD eess.AS 81%

How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes

Mahnoor Fatima Saad, Ziad Al-Halah

机构 * University of Utah(犹他大学)

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract);分类 cs.CV、eess.AS

Comments Accepted to ICCV 2025. Project Page: https://mahnoor-fatima-saad.github.io/m-capa.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00265 2025-08-06 cs.CV 80%

Multimodal Referring Segmentation: A Survey

Henghui Ding, Song Tang, Shuting He, Chang Liu, Zuxuan Wu, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai University of Finance and Economics(上海财经大学) ByteDance Inc(字节跳动公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://github.com/henghuiding/Awesome-Multimodal-Referring-Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08328 2025-08-06 eess.AS 74%

Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation

Tsun-An Hsieh, Heeyoul Choi, Minje Kim

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Journal ref Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03166 2025-08-06 cs.SD cs.LG eess.AS 74%

MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction

Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov

专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS

Comments 5 pages, 2 figures, 1 table. Accepted for presentation at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21254 2025-08-06 cs.CV cs.AI cs.MM cs.SD eess.AS 70%

Vision-to-Music Generation: A Survey

Zhaokai Wang, Chenxi Bao, Le Zhuo, Jingrui Han, Yang Yue, Yihong Tang, Victor Shea-Jay Huang, Yue Liao

机构 * Shanghai Jiao Tong University(上海交通大学) Music Tech Lab, DynamiX(音乐科技实验室,DynamiX) Shanghai AI Laboratory(上海人工智能实验室) Beijing Film Academy(北京电影学院) Tsinghua University(清华大学) McGill University(麦吉尔大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Journal ref ISMIR 2025 "A Survey on Vision to Music Generation: Methods, Datasets, Evaluation, and Challenges"

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22053 2025-08-06 cs.SD cs.MA cs.MM eess.AS 62%

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tencent AI Lab(腾讯人工智能实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2508.03694 2025-08-06 cs.CV 79%

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Jianfeng Feng, Chenyang Si, Yanwei Fu, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) Fudan University(复旦大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) NVIDIA(NVIDIA公司) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Project page: https://vchitect.github.io/LongVie-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07454 2025-08-06 cs.CV 70%

How Can Objects Help Video-Language Understanding?

Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo, Chen Sun

机构 * Brown University(布朗大学) Samsung Electronics(三星电子)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03009 2025-08-06 cs.CV cs.AI 62%

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin, Lin Liu, Hongbo Xu, Yongwei Nie, Fei Yu, Fei Ma

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18688 2025-08-06 cs.CV cs.AI 62%

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

Faraz Waseem, Muhammad Shahzad

机构 * University Of Reading(阅读大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 35 pages, 18 figures, Manuscript submitted to ACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03651 2025-08-06 cs.HC cs.AI 57%

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Anhong Guo

机构 * University of Michigan(密歇根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments ACM ASSETS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03518 2025-08-06 cs.IR cs.LG 50%

Parameter-Efficient Single Collaborative Branch for Recommendation

Marta Moscati, Shah Nawaz, Markus Schedl

机构 * Institute of Computational Perception, Johannes Kepler University Linz(计算感知研究所,林茨约瑟夫·施密特大学) AI Lab, Linz Institute of Technology(林茨技术学院人工智能实验室)

专题命中 视频多模态 :multimodal(abstract)

Comments 5 pages

Journal ref Proceedings of the Nineteenth ACM Conference on Recommender Systems (RecSys'25), September 22-26, 2025, Prague, Czech Republic. ACM, New York, NY, USA

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2508.03494 2025-08-06 cs.CV 79%

Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval

Shreyank N Gowda, Xiaobo Jin, Christian Wagner

机构 * School of Computer Science, The University of Nottingham, NG8 1BB Nottingham, U.K.(计算机科学学院,诺丁汉大学) Department of Intelligent Science, Xi’an Jiaotong-Liverpool University, China, 215123.(智能科学系,西安交通大学利物浦大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03298 2025-08-06 cs.SE 78%

GUI-ReRank: Enhancing GUI Retrieval with Multi-Modal LLM-based Reranking

Kristian Kolthoff, Felix Kretzer, Christian Bartelt, Alexander Maedche, Simone Paolo Ponzetto

专题命中 跨模态检索 :multi-modal(title);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03091 2025-08-06 cs.AI cs.CR cs.CV 73%

T2UE: Generating Unlearnable Examples from Text Descriptions

Xingjun Ma, Hanxun Huang, Tianwei Song, Ye Sun, Yifeng Gao, Yu-Gang Jiang

机构 * Fudan University(复旦大学) The University of Melbourne(墨尔本大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments To appear in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03644 2025-08-06 cs.CL cs.CV cs.IR 62%

Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?

Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen, Junjie Yang, Yao Wan, Weiwei Lin

机构 * South China University of Technology(华南理工大学) Huazhong University of Science and Technology(华中科技大学) University of Maryland(马里兰大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments In submission. Project website: https://double-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03262 2025-08-06 cs.CL cs.AI 62%

Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?

Junhyuk Choi, Hyeonchu Park, Haemin Lee, Hyebeen Shin, Hyun Joung Jin, Bugeun Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2411.04954 2025-08-06 cs.CV 84%

CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM

Jingwei Xu, Chenyu Wang, Zibo Zhao, Wen Liu, Yi Ma, Shenghua Gao

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Transcengram DeepSeek AI University of Hong Kong(香港大学)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments Project page: https://cad-mllm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03426 2025-08-06 cs.CV cs.AI cs.LG 81%

R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

Futian Wang, Yuhan Qiao, Xiao Wang, Fuling Wang, Yuxiang Zhang, Dengdi Sun

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21893 2025-08-06 cs.CV 79%

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

Saeed Ghorbani

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03690 2025-08-06 cs.CV cs.RO 57%

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

Youquan Liu, Lingdong Kong, Weidong Yang, Ao Liang, Jianxiong Gao, Yang Wu, Xiang Xu, Xin Li, Linfeng Li, Runnan Chen, Ben Fei

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Preprint; 10 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03669 2025-08-06 cs.CV cs.RO 57%

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

Katherine Liu, Sergey Zakharov, Dian Chen, Takuya Ikeda, Greg Shakhnarovich, Adrien Gaidon, Rares Ambrus

机构 * Toyota Research Institute(丰田研究院) Woven by Toyota(丰田编织) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures. This version has typo fixes on top of the version published at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03539 2025-08-06 cs.CV 57%

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

Long Qian, Bingke Zhu, Yingying Chen, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(自动化研究所基础模型研究中心,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Objecteye Inc.(Objecteye公司)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03535 2025-08-06 cs.CV 57%

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu, Wenshuo Chen, Yutao Yue

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03320 2025-08-06 cs.CV 57%

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Peiyu Wang, Yi Peng, Yimeng Gan, Liang Hu, Tianyidan Xie, Xiaokun Wang, Yichen Wei, Chuanxin Tang, Bo Zhu, Changshi Li, Hongyang Wei, Eric Li, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Multimodality Team, Skywork AI(Skywork AI 多模态团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏