arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-01 至 2025-10-01 共收录 77 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 5 篇

2509.26378 2025-10-01 cs.IR cs.CV 79%

MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval

Junjie Zhou, Ze Liu, Lei Xiong, Jin-Ge Yao, Yueze Wang, Shitao Xiao, Fenfen Lin, Miguel Hu Chen, Zhicheng Dou, Siqi Bao, Defu Lian, Yongping Xiong, Zheng Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26330 2025-10-01 cs.CV cs.IR 70%

SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval

Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang

机构 * National Sun Yat-sen University(国立中山大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26012 2025-10-01 cs.CV 57%

SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval

Yuqi Xiao, Yingying Zhu

机构 * Yuqi Xiao, Yingying Zhu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 10 篇

2509.24361 2025-10-01 cs.CV cs.AI cs.HC 84%

UI-UG: A Unified MLLM for UI Understanding and Generation

Hao Yang, Weijie Qiu, Ru Zhang, Zhou Fang, Ruichao Mao, Xiaoyu Lin, Maji Huang, Zhaosong Huang, Teng Guo, Shuoyang Liu, Hai Rao

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26231 2025-10-01 cs.CV 83%

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

Jiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang, Kai Wang, Gao Huang, Humphrey Shi

机构 * Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25681 2025-10-01 cs.RO cs.CV 83%

dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

Junjie Wen, Minjie Zhu, Jiaming Liu, Zhiyuan Liu, Yicun Yang, Linfeng Zhang, Shanghang Zhang, Yichen Zhu, Yi Xu

机构 * Midea Group(美的集团) Peking University(北京大学) Shanghai Jiaotong University(上海交通大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments technique report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26644 2025-10-01 cs.CV cs.AI cs.LG 81%

Stitch: Training-Free Position Control in Multimodal Diffusion Transformers

Jessica Bader, Mateusz Pach, Maria A. Bravo, Serge Belongie, Zeynep Akata

机构 * Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(海德堡-慕尼黑研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Copenhagen(哥本哈根大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26641 2025-10-01 cs.CV 79%

Query-Kontext: An Unified Multimodal Model for Image Generation and Editing

Yuxin Song, Wenkai Dong, Shizun Wang, Qi Zhang, Song Xue, Tao Yuan, Hu Yang, Haocheng Feng, Hang Zhou, Xinyan Xiao, Jingdong Wang

机构 * Baidu VIS(百度视觉) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 62%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25401 2025-10-01 cs.LG cs.AI cs.PF 57%

FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers

Liang Qiao, Yue Dai, Yeqi Huang, Hongyu Kan, Jun Shi, Hong An

机构 * University of Science and Technology of China(中国科学技术大学) University of Edinburgh(爱丁堡大学) University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25229 2025-10-01 cs.AI 57%

Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models

Lukas Petersson, Axel Backlund, Axel Wennstöm, Hanna Petersson, Callum Sharrock, Arash Dabiri

机构 * Andon Labs(安登实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 8 figures, submitted for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06266 2025-10-01 cs.RO 50%

ADPro: a Test-time Adaptive Diffusion Policy via Manifold-constrained Denoising and Task-aware Initialization for Robotic Manipulation

Zezeng Li, Rui Yang, Ruochen Chen, ZhongXuan Luo, Liming Chen

机构 * École Centrale de Lyon(里昂中央理工大学) Dalian University of Technology(大连理工大学)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02353 2025-10-01 cs.RO cs.LG 50%

Controllable Motion Generation via Diffusion Modal Coupling

Luobin Wang, Hongzhan Yu, Chenning Yu, Sicun Gao, Henrik Christensen

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 16 篇

2509.24297 2025-10-01 cs.CL cs.AI 81%

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen, Yalun Wu, Yingji Liang, Yijin Guo, Farong Wen, Wenzhe Li, Xuezhi Zhao, Qi Jia, Guangtao Zhai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26061 2025-10-01 eess.IV cs.CV 79%

Multi-modal Liver Segmentation and Fibrosis Staging Using Real-world MRI Images

Yang Zhou, Kunhao Yuan, Ye Wei, Jishizhan Chen

机构 * Multiscale X-ray Imaging (MXI) Lab, Department of Mechanical Engineering, University College London(多尺度X射线成像实验室,机械工程系,伦敦大学学院) Centre for Clinical Brain Sciences, University of Edinburgh(临床脑科学中心,爱丁堡大学) MRC Weatherall Institute of Molecular Medicine, University of Oxford(MRC韦伯尔分子医学研究所,牛津大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25946 2025-10-01 cs.AI 79%

Automated Model Discovery via Multi-modal & Multi-step Pipeline

Lee Jung-Mok, Nam Hyeon-Woo, Moon Ye-Bin, Junhyun Nam, Tae-Hyun Oh

机构 * Dept. of Electrical Engineering, POSTECH(POSTECH电子工程系) Samsung Electronics(三星电子) School of Computing, KAIST(KAIST计算机科学学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25564 2025-10-01 cs.CV 79%

FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology

Faizan Farooq Khan, Yousef Radwan, Eslam Abdelrahman, Abdulwahab Felemban, Aymen Mir, Nico K. Michiels, Andrew J. Temple, Michael L. Berumen, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(国王阿卜杜勒-阿齐兹大学科学与技术学院) Red Sea Research Center, KAUST(红海研究中心,KAUST) Tübingen University(图宾根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 3 figures 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25559 2025-10-01 cs.AI cs.LG 79%

Radiology's Last Exam (RadLE): Benchmarking Frontier Multimodal AI Against Human Experts and a Taxonomy of Visual Reasoning Errors in Radiology

Suvrankar Datta, Divya Buchireddygari, Lakshmi Vennela Chowdary Kaza, Mrudula Bhalke, Kautik Singh, Ayush Pandey, Sonit Sai Vasipalli, Upasana Karnwal, Hakikat Bir Singh Bhatti, Bhavya Ratan Maroo, Sanjana Hebbar, Rahul Joseph, Gurkawal Kaur, Devyani Singh, Akhil V, Dheeksha Devasya Shama Prasad, Nishtha Mahajan, Ayinaparthi Arisha, Rajesh Vanagundi, Reet Nandy, Kartik Vuthoo, Snigdhaa Rajvanshi, Nikhileswar Kondaveeti, Suyash Gunjal, Rishabh Jain, Rajat Jain, Anurag Agrawal

机构 * Centre for Responsible Autonomous Systems in Healthcare (CRASH) Lab, Koita Centre for Digital Health(负责任的自主医疗系统中心(CRASH)实验室、Koita数字健康中心) Ashoka University(阿什oka大学) Independent Researcher(独立研究者) Koita Centre for Digital Health(Koita数字健康中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 29 pages, 7 figures, 7 tables, includes Annexure (1). Part of the work accepted at RSNA 2025 (Cutting Edge Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06637 2025-10-01 cs.MM 79%

SCI-Reason: A Dataset with Chain-of-Thought Rationales for Complex Multimodal Reasoning in Academic Areas

Chenghao Ma, Haihong E., Junpeng Ding, Jun Zhang, Ziyan Ma, Huang Qing, Bofei Gao, Liang Chen, Yifan Zhu, Meina Song

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

Comments Submitted to ICCV 2025. 11 pages (including references)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26047 2025-10-01 cs.CV 70%

DGM4+: Dataset Extension for Global Scene Inconsistency

Gagandeep Singh, Samudi Amarsinghe, Priyanka Singh, Xue Li

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25502 2025-10-01 cs.CV 70%

Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection

Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Junyan Ye, Ke-Yue Zhang, Yue Zhou, Peng Jin, Bin Li, Taiping Yao, Shouhong Ding

机构 * Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen Key Laboratory of Media Security, and SZU-AFS Joint Innovation Center for AI Technology(广东省智能信息处理重点实验室、深圳媒体安全重点实验室及深圳大学-AFS联合人工智能技术创新中心) Tencent Youtu Lab(腾讯优图实验室) Peking University(北京大学) Sun Yat-Sen University(中山大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25818 2025-10-01 cs.CV cs.AI cs.CL 67%

VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki, Komei Sugiura

机构 * Keio University(庆应大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03214 2025-10-01 cs.CL cs.AI cs.CV 67%

iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs

Julius Mayer, Mohamad Ballout, Serwan Jassim, Farbod Nosrat Nezami, Elia Bruni

机构 * Institute of Cognitive Science, Osnabrück University(认知科学研究所,奥斯纳布吕克大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21611 2025-10-01 cs.CL cs.AI cs.LG 62%

When Does Multimodality Lead to Better Time Series Forecasting?

Xiyuan Zhang, Boran Han, Haoyang Fang, Abdul Fatir Ansari, Shuai Zhang, Danielle C. Maddix, Cuixiong Hu, Andrew Gordon Wilson, Michael W. Mahoney, Hao Wang, Yan Liu, Huzefa Rangwala, George Karypis, Bernie Wang

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26639 2025-10-01 cs.CV cs.RO 57%

Benchmarking Egocentric Visual-Inertial SLAM at City Scale

Anusha Krishnan, Shaohui Liu, Paul-Edouard Sarlin, Oscar Gentilhomme, David Caruso, Maurizio Monge, Richard Newcombe, Jakob Engel, Marc Pollefeys

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Meta Reality Labs Research(Meta现实实验室) Microsoft Spatial AI Lab(微软空间人工智能实验室)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26500 2025-10-01 eess.SP cs.AI cs.NI 57%

Indoor/Outdoor Spectrum Sharing Enabled by GNSS-based Classifiers

Hossein Nasiri, Muhammad Iqbal Rochman, Monisha Ghosh

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Comments To be published in the proceedings of IEEE Military Communications Conference (MILCOM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25811 2025-10-01 cs.CV cs.LG 57%

Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition

Zichen Liang, Jingjing Fei, Jie Wang, Zheming Yang, Changqing Li, Pei Wu, Minghui Qiu, Fei Yang, Xialei Liu

机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26221 2025-10-01 cs.LG 50%

Marginal Flow: a flexible and efficient framework for density estimation

Marcello Massimo Negri, Jonathan Aellen, Manuel Jahn, AmirEhsan Khorashadizadeh, Volker Roth

机构 * Department of Computer Science University of Basel(计算机科学系伯尔尼大学)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.08838 2025-10-01 cs.LG 50%

AQuaMaM: An Autoregressive, Quaternion Manifold Model for Rapidly Estimating Complex SO(3) Distributions

Michael A. Alcorn

机构 * USDA(美国农业部)

专题命中 多模态评测 :multimodal(abstract)

Comments Accepted to the Non-Euclidean Foundation Models and Geometric Learning workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态Agent 2 篇

2509.26161 2025-10-01 cs.AI cs.SE 79%

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态Agent :MLLM(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏