arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-29 至 2025-09-29 共收录 79 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 7 篇

2509.21356 2025-09-29 cs.CV cs.AI 62%

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood

机构 * Rensselaer Polytechnic Institute, NY, USA(罗文学院) IBM Research, Almaden, CA, USA(IBM研究院) Stanford University, CA, USA(斯坦福大学) Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments In proceedings MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21336 2025-09-29 cs.IR cs.CL 57%

HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores

Guohang Yan, Yue Zhang, Pinlong Cai, Ding Wang, Song Mao, Hongwei Zhang, Yaoze Zhang, Hairong Zhang, Xinyu Cai, Botian Shi

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21507 2025-09-29 cs.CE 50%

QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance

Haoxue Wang, Keli Wen, Yuante Li, Qiancheng Qu, Xiangxu Mu, Xinjie Shen, Jiaqi Gao, Chenyang Chang, Chuhan Xie, San Yu Cheung, Zhuoyuan Hu, Xinyu Wang, Sirui Bi, Bi'an Du

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 12 篇

2509.21360 2025-09-29 cs.CV cs.AI 81%

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22570 2025-09-29 cs.AI 79%

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

Qi Mao, Tinghan Yang, Jiahao Li, Bin Li, Libiao Jin, Yan Lu

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室) Communication University of China(中国传媒大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 78%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03255 2025-09-29 cs.CV 70%

DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation

Qingdong He, Jinlong Peng, Pengcheng Xu, Boyuan Jiang, Xiaobin Hu, Donghao Luo, Yong Liu, Yabiao Wang, Chengjie Wang, Xiangtai Li, Jiangning Zhang

机构 * Youtu Lab, Tencent(腾讯优图实验室) Western University(西部大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 62%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21375 2025-09-29 cs.CV cs.AI 62%

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

Aleksa Jelaca, Ying Jiao, Chang Tian, Marie-Francine Moens

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments text-to-image generation, automatic prompt, DPO, Counterfactual

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22573 2025-09-29 cs.RO cs.CV 57%

MINT-RVAE: Multi-Cues Intention Prediction of Human-Robot Interaction using Human Pose and Emotion Information from RGB-only Camera Data

Farida Mohsen, Ali Safa

机构 * College of Science and Engineering, Hamad Bin Khalifa University(哈马德·本·哈利法大学科学与工程学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22485 2025-09-29 cs.CV 57%

Group Critical-token Policy Optimization for Autoregressive Image Generation

Guohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang, Yaning Pan, Mingde Yao, Jie Xiao, Linjiang Huang, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) CUHK(香港中文大学) Beihang University(北京航空航天大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Code is available at https://github.com/zghhui/GCPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21997 2025-09-29 cs.CV 57%

Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors

Youxu Shi, Suorong Yang, Dong Liu

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21760 2025-09-29 cs.CV 57%

UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models

Lan Chen, Yuchao Gu, Qi Mao

机构 * MIPG, Communication University of China(信息与通信大学) Show Lab, National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20253 2025-09-29 cs.RO cs.AI 57%

AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving

Jinhao Chai, Anqing Jiang, Hao Jiang, Shiyi Mu, Zichong Gu, Hao Sun, Shugong Xu

机构 * School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China(信息工程学院,上海大学) Bosch Corporate Research, Bosch (China) Investment Ltd., Shanghai, China(博世企业研究,博世(中国)投资有限公司) School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China(机械工程学院,上海交通大学) Xi'an Jiaotong-Liverpool University, Suzhou, China(西安交通大学利物浦大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21664 2025-09-29 cs.RO cs.LG 50%

Generating Stable Placements via Physics-guided Diffusion Models

Philippe Nadeau, Miguel Rogel, Ivan Bilić, Ivan Petrović, Jonathan Kelly

机构 * STARS Laboratory, University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究所STARS实验室) Laboratory for Autonomous Systems and Mobile Robotics(自主系统与移动机器人实验室) University of Zagreb Faculty of Electrical Engineering and Computing(Zagreb大学电气工程与计算学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments Submitted to the IEEE International Conference on Robotics and Automation 2026, Vienna, Austria, June 1-5, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 20 篇

2509.21451 2025-09-29 cs.CV cs.CL 84%

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

Abdul Waheed, Zhen Wu, Dareen Alharthi, Seungone Kim, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21365 2025-09-29 cs.CV cs.AI 84%

MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation

Zhicheng Du, Qingyang Shi, Jiasheng Lu, Yingshan Liang, Xinyu Zhang, Yiran Wang, Peiwu Qin

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05271 2025-09-29 cs.CV 83%

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yiming Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jiaye Ge, Kai Chen, Kaipeng Zhang, Limin Wang, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) Nanjing University(南京大学) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22377 2025-09-29 cs.CV 79%

Effectiveness of Large Multimodal Models in Detecting Disinformation: Experimental Results

Yasmina Kheddache, Marc Lalonde

机构 * Département d’informatique et recherche opérationnelle (D.I.R.O.) Université de Montréal Pavillon André-Aisenstadt 2920, chemin de la Tour Montréal (QC) H3T 1N8(蒙特利尔大学计算机与运筹学系) R&D Dept. Computer Research Institute of Montreal (CRIM) 405 Ogilvy Ave., #101 Montreal, Qc, Canada(蒙特利尔计算机研究 institute)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22019 2025-09-29 cs.CV 79%

EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking

Yuki Sakai, Ryosuke Furuta, Juichun Yen, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted to the I-HFM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21805 2025-09-29 cs.CL 79%

Towards Minimal Causal Representations for Human Multimodal Language Understanding

Menghua Jiang, Yuncheng Jiang, Haifeng Hu, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院) School of Electronics and Information Technology, Sun Yat-sen University(中山大学电子与信息学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21600 2025-09-29 cs.AI cs.LG 79%

Automated and Interpretable Survival Analysis from Multimodal Data

Mafalda Malafaia, Peter A. N. Bosman, Coen Rasch, Tanja Alderliesten

机构 * Centrum Wiskunde & Informatica(数学与信息学研究中心) Delft University of Technology(代尔夫特理工大学) Leiden University Medical Center(莱顿大学医学中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 4 figures; 4 tables; 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16044 2025-09-29 eess.AS cs.LG eess.IV eess.SP 79%

Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation

Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Espy-Wilson

机构 * Department of Electrical \& Computer Engineering Institute for Advanced Computer Studies School of Medicine

专题命中 多模态评测 :multimodal(title,abstract);分类 eess.AS

Comments Accepted to be presented at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21352 2025-09-29 cs.CV cs.LG 79%

Improving Autism Detection with Multimodal Behavioral Analysis

William Saakyan, Matthias Norden, Lola Eversmann, Simon Kirsch, Muyu Lin, Simon Guendelman, Isabel Dziobek, Hanna Drimalla

机构 * Center for Cognitive Interaction Technology (CITEC), Bielefeld University(认知交互技术中心(CITEC),比勒菲尔德大学) Institute of Psychology, Humboldt University of Berlin(心理学研究所,洪堡大学) Department of Psychiatry and Psychotherapy, Medical Center-University of Freiburg(精神病学与心理治疗系,弗赖堡医学院-大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21396 2025-09-29 cs.CV cs.LG 74%

mmHSense: Multi-Modal and Distributed mmWave ISAC Datasets for Human Sensing

Nabeel Nisar Bhat, Maksim Karnaukh, Stein Vandenbroeke, Wouter Lemoine, Jakob Struye, Jesus Omar Lacruz, Siddhartha Kumar, Mohammad Hossein Moghaddam, Joerg Widmer, Rafael Berkvens, Jeroen Famaey

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22518 2025-09-29 cs.AI cs.LG 57%

REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model

Bo Li, Guanzhi Deng, Ronghao Chen, Junrong Yue, Shuo Zhang, Qinghua Zhao, Linqi Song, Lijie Wen

机构 * Tsinghua University(清华大学) City University of Hong Kong(香港城市大学) Peking University(北京大学) Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航) Baidu Inc(百度公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22228 2025-09-29 cs.CV 57%

UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

Jun He, Yi Lin, Zilong Huang, Jiacong Yin, Junyan Ye, Yuchuan Zhou, Weijia Li, Xiang Zhang

机构 * Sun Yat-sen University(中山大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21450 2025-09-29 cs.CL 57%

LLM-Based Support for Diabetes Diagnosis: Opportunities, Scenarios, and Challenges with GPT-5

Gaurav Kumar Gupta, Nirajan Acharya, Pranal Pande

机构 * Youngstown State University(扬斯敦州立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01188 2025-09-29 cs.LG cs.AI 57%

SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy

Zhuo Yang, Jiaqing Xie, Shuaike Shen, Daolang Wang, Yeyun Chen, Ben Gao, Shuzhou Sun, Biqing Qi, Dongzhan Zhou, Lei Bai, Linjiang Chen, Shufei Zhang, Qinying Gu, Jun Jiang, Tianfan Fu, Yuqiang Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Xidian University(西安电子科技大学) Nanjing University(南京大学) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) North University of China(北方大学) Center for Machine Vision and Signal Analysis (CMVS), University of Oulu(奥卢大学机器视觉与信号分析中心) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) Shanghai Innovation Institute(上海创新研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21875 2025-09-29 cs.CL 57%

WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild

Linhao Zhang, Jian Zhang, Bokai Lei, Chuhan Wu, Aiwei Liu, Wei Jia, Xiao Zhou

机构 * Pattern Recognition Center, WeChat AI, Tencent Inc(模式识别中心、微信AI、腾讯公司)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏