arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-22 至 2025-09-22 共收录 63 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2509.15246 2025-09-22 cs.GR cs.AI 79%

GenCAD-3D: CAD Program Generation using Multimodal Latent Space Alignment and Synthetic Dataset Balancing

Nomi Yu, Md Ferdous Alam, A. John Hart, Faez Ahmed

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 9 figures, 15 pages. Accepted and soon published in the ASME Journal of Mechanical Design

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14476 2025-09-22 cs.CV cs.AI cs.MM 67%

AToken: A Unified Tokenizer for Vision

Jiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn, Yanjun Wang, Chen Chen, Afshin Dehghan, Yinfei Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 30 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 62%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15270 2025-09-22 cs.CV cs.AI 62%

PRISM: Phase-enhanced Radial-based Image Signature Mapping framework for fingerprinting AI-generated images

Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, Roberto Di Pietro

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 10 篇

2509.15241 2025-09-22 cs.CV cs.CL 86%

M-PACE: Mother Child Framework for Multimodal Compliance

Shreyash Verma, Amit Kesari, Vinayak Trivedi, Anupam Purwar, Ratnesh Jamidar

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments The M-PACE framework uses a "mother-child" AI model system to automate and unify compliance checks for ads, reducing costs while maintaining high accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15361 2025-09-22 cs.CL cs.AI cs.MM 82%

Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing

Zichen Wu, Hsiu-Yuan Huang, Yunfang Wu

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) MOE Key Laboratory of Computational Linguistics, Peking University(北京大学教育部计算语言学重点实验室) National Key Laboratory for Multimedia Information Processing, Peking University(北京大学国家多媒体信息处理重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15701 2025-09-22 cs.CL cs.SD eess.AS 81%

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

Ke Wang, Wenning Wei, Yan Deng, Lei He, Sheng Zhao

机构 * Microsoft, Beijing, China(微软北京研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15839 2025-09-22 cs.CL 79%

Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems

Zhongze Luo, Zhenshuai Yin, Yongxin Guo, Zhichao Wang, Jionghao Zhu, Xiaoying Tang

机构 * School of Science(科学学院) Engineering, The Chinese University of Hong Kong, Shenzhen, China(工程学院,香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13212 2025-09-22 cs.CV 79%

Semantic Change Detection of Roads and Bridges: A Fine-grained Dataset and Multimodal Frequency-driven Detector

Qingling Shu, Sibao Chen, Xiao Wang, Zhihui You, Wei Lu, Jin Tang, Bin Luo

机构 * MOE Key Lab of ICSP, IMIS Lab of Anhui Province, Anhui Provincial Key Lab of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University, Hefei, China(教育部ICSP重点实验室、安徽IMIS实验室、安徽省多模态认知计算重点实验室、计算机科学与技术学院、安徽大学,合肥,中国) School of Public Safety and Emergency Management, Anhui University of Science and Technology, Hefei, China(公共安全与应急管理学院、安徽理工大学,合肥,中国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15706 2025-09-22 cs.CV cs.AI physics.ao-ph 62%

SGMAGNet: A Baseline Model for 3D Cloud Phase Structure Reconstruction on a New Passive Active Satellite Benchmark

Chi Yang, Fu Wang, Xiaofei Yang, Hao Huang, Weijia Cao, Xiaowen Chu

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15662 2025-09-22 cs.MM cs.SD eess.AS 62%

Jamendo-QA: A Large-Scale Music Question Answering Dataset

Junyoung Koh, Soo Yong Kim, Yongwon Choi, Gyu Hyeong Choi

机构 * Department of Artificial Intelligence, Yonsei University(延世大学人工智能系) MAAP LAB, MODULABS(MODULABS MAAP实验室) KRAFTON AI Matics Department of Media Software, Sungkyul University(顺天大学媒体软件系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.MM、eess.AS

Comments 4 pages, 8 figures. Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07894 2025-09-22 cs.AI 57%

HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?

Fangchen Yu, Haiyuan Wan, Qianjia Cheng, Yuchen Zhang, Jiacheng Chen, Fujun Han, Yulun Wu, Junchi Yao, Ruilizhen Hu, Ning Ding, Yu Cheng, Tao Chen, Lei Bai, Dongzhan Zhou, Yun Luo, Ganqu Cui, Peng Ye

机构 * Shanghai AI Laboratory(上海人工智能实验室) CUHK-Shenzhen(香港中文大学(深圳)) CUHK(香港中文大学) Tsinghua University(清华大学) Zhejiang University(浙江大学) Peking University(北京大学) University of Electronic Science and Technology of China(电子科技大学) Fudan University(复旦大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17464 2025-09-22 cs.CL 57%

HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning

Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang

机构 * University of New South Wales(新南威尔士大学) Data61, CSIRO(Data61,CSIRO)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CL

Comments Accepted by EMNLP2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09613 2025-09-22 cs.SI cs.CY 50%

How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis

Herun Wan, Minnan Luo, Zihan Ma, Guang Dai, Xiang Zhao

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 5 篇

2508.16051 2025-09-22 cs.AI 79%

MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs

Yiheng Hu, Xiaoyang Wang, Qing Liu, Xiwei Xu, Qian Fu, Wenjie Zhang, Liming Zhu

机构 * University of New South Wales(新南威尔士大学) CSIRO Data61(CSIRO数据61)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20474 2025-09-22 q-fin.TR cs.CL cs.LG 79%

MountainLion: A Multi-Modal LLM-Based Agent System for Interpretable and Adaptive Financial Trading

Siyi Wu, Junqiao Wang, Zhaoyang Guan, Leyi Zhao, Xinyuan Song, Xinyu Ying, Dexu Yu, Jinhao Wang, Hanlin Zhang, Michele Pak, Yangfan He, Yi Xin, Jianhui Wang, Tianyu Shi

机构 * The University of Texas at Arlington(德克萨斯大学阿灵顿分校) Sichuan University(四川大学) Northwestern University(西北大学) Indiana University(印第安纳大学) Emory University(埃默里大学) Nankai University(南开大学) MountainLion Research(MountainLion研究机构) Xi’an University of Electronic Science and Technology(西安电子科技大学) Kyoto University(京都大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Nanjing university(南京大学) Tsinghua University(清华大学) University of Toronto(多伦多大学)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03603 2025-09-22 cs.AI cs.LG 79%

Towards deployment-centric multimodal AI beyond vision and language

Xianyuan Liu, Jiayang Zhang, Shuo Zhou, Thijs L. van der Plas, Avish Vijayaraghavan, Anastasiia Grishina, Mengdie Zhuang, Daniel Schofield, Christopher Tomlinson, Yuhan Wang, Ruizhe Li, Louisa van Zeeland, Sina Tabakhi, Cyndie Demeocq, Xiang Li, Arunav Das, Orlando Timmerman, Thomas Baldwin-McDonald, Jinge Wu, Peizhen Bai, Zahraa Al Sahili, Omnia Alwazzan, Thao N. Do, Mohammod N. I. Suvon, Angeline Wang, Lucia Cipolina-Kun, Luigi A. Moretti, Lucas Farndale, Nitisha Jain, Natalia Efremova, Yan Ge, Marta Varela, Hak-Keung Lam, Oya Celiktutan, Ben R. Evans, Alejandro Coca-Castro, Honghan Wu, Zahraa S. Abdallah, Chen Chen, Valentin Danchev, Nataliya Tkachenko, Lei Lu, Tingting Zhu, Gregory G. Slabaugh, Roger K. Moore, William K. Cheung, Peter H. Charlton, Haiping Lu

机构 * Centre for Machine Intelligence, University of Sheffield, Sheffield, UK(机器智能中心,谢菲尔德大学) School of Computer Science, University of Sheffield, Sheffield, UK(计算机科学学院,谢菲尔德大学) The Alan Turing Institute, London, UK(艾伦·图灵研究所,伦敦) Department of Metabolism, Digestion and Reproduction, Imperial College London, London, UK(代谢、消化与生殖部门,伦敦帝国学院) Department of Applied AI, Simula Research Laboratory, Oslo, Norway(应用人工智能部门,Simula研究实验室,奥斯陆) Information School, University of Sheffield, Sheffield, UK(信息学院,谢菲尔德大学) NHS England, Leeds, UK(英国英格兰国家医疗服务体系,利兹) Institute of Health Informatics, University College London, London, UK(健康信息研究所,伦敦大学学院) Department of Engineering, King’s College London, London, UK(工程部门,伦敦国王学院) Department of Computing Science, University of Aberdeen, Aberdeen, UK(计算科学部门,阿伯丁大学) School of Informatics, University of Edinburgh, Edinburgh, UK(信息学院,爱丁堡大学) School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK(工程数学与技术学院,布里斯托尔大学) Department of Informatics, King’s College London, London, UK(信息部门,伦敦国王学院) Department of Earth Sciences, University of Cambridge, Cambridge, UK(地球科学部门,剑桥大学) Department of Computer Science, University of Manchester, Manchester, UK(计算机科学部门,曼彻斯特大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15635 2025-09-22 cs.AI 70%

MicroRCA-Agent: Microservice Root Cause Analysis Method Based on Large Language Model Agents

Pan Tang, Shixiang Tang, Huanqi Pu, Zhiqing Miao, Zhixing Wang

机构 * School of Communication and Information Engineering, Shanghai University, Shanghai, China(上海大学通信与信息工程学院) School of Communication and Electronic Engineering, East China Normal University, Shanghai, China(华东师范大学通信与电子工程学院) School of Information and Electronics, Beijing Institute of Technology, Beijing, China(北京理工大学信息与电子学院)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20148 2025-09-22 cs.AI cs.HC cs.MA 57%

The Anatomy of a Personal Health Agent

A. Ali Heydari, Ken Gu, Vidya Srinivas, Hong Yu, Zhihan Zhang, Yuwei Zhang, Akshay Paruchuri, Qian He, Hamid Palangi, Nova Hammerquist, Ahmed A. Metwally, Brent Winslow, Yubin Kim, Kumar Ayush, Yuzhe Yang, Girish Narayanswamy, Maxwell A. Xu, Jake Garrison, Amy Armento Lee, Jenny Vafeiadou, Ben Graef, Isaac R. Galatzer-Levy, Erik Schenck, Andrew Barakat, Javier Perez, Jacqueline Shreibati, John Hernandez, Anthony Z. Faranesh, Javier L. Prieto, Connor Heneghan, Yun Liu, Jiening Zhan, Mark Malhotra, Shwetak Patel, Tim Althoff, Xin Liu, Daniel McDuff, Xuhai "Orson" Xu

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Minor updates to the manuscript (V2)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 11 篇

2509.15578 2025-09-22 cs.CV cs.AI 84%

Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion

Shanghong Li, Chiam Wen Qi Ruth, Hong Xu, Fang Liu

机构 * Singapore University of Social Sciences(新加坡社会科学研究大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16149 2025-09-22 cs.CV 83%

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu, Jiahui Gao, Jipeng Zhang, Xiaofang Zhou

机构 * HKUST(香港科技大学) HKU(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16017 2025-09-22 cs.CV 83%

DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching

Meng Yang, Fan Fan, Zizhuo Li, Songchu Deng, Yong Ma, Jiayi Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15852 2025-09-22 cs.MM 79%

Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction

Yueheng Jiang, Peng Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15277 2025-09-22 cs.MM cs.LG 79%

Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction

Qin Chao, Eunsoo Kim, Boyang Li

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡) Business School, University of Seoul, Republic of Korea(商学学院,首尔大学,韩国) Alibaba Group and the Alibaba-NTU Joint Research Institute, Singapore(阿里巴巴集团及阿里巴巴-南洋理工大学联合研究机构,新加坡)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19668 2025-09-22 eess.SP cs.AI cs.CL cs.LG 76%

SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning

Mingsheng Cai, Jiuming Jiang, Wenhao Huang, Che Liu, Rossella Arcucci

机构 * The University of Edinburgh(爱丁堡大学) Imperial College London(帝国理工学院) Shenzhen Yinwang Intelligent Technology Co., Ltd(深圳英伟达智能技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

Comments Findings of The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 73%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 73%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16170 2025-09-22 cs.CV 70%

UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation

Xiaoqi Zhao, Youwei Pang, Chenyang Yu, Lihe Zhang, Huchuan Lu, Shijian Lu, Georges El Fakhri, Xiaofeng Liu

机构 * Yale University, USA(耶鲁大学) Nanyang Technological University, Singapore(南洋理工大学) Dalian University of Technology, China(大连理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15990 2025-09-22 cs.CV 57%

DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis

Jérémie Stym-Popper, Nathan Painchaud, Clément Rambour, Pierre-Yves Courand, Nicolas Thome, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 9 pages, Accepted at MIDL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18353 2025-09-22 cs.LG 50%

A Data-Driven Review of Remote Sensing-Based Data Fusion in Precision Agriculture from Foundational to Transformer-Based Techniques

Mahdi Saki, Rasool Keshavarz, Daniel Franklin, Mehran Abolhasan, Justin Lipman, Negin Shariati

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 22 pages, 13 figures, 3 tables, Journal

详情

展开后加载摘要…

URL PDF HTML 收藏