arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 13 篇

2406.10424 2025-09-16 cs.CV cs.AI 84%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 81%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11082 2025-09-16 cs.CV cs.RO 79%

Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation

Zongwu Xie, Kaijie Yun, Yang Liu, Yiming Ji, Han Li

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10704 2025-09-16 cs.AI cs.CV 73%

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

Xingchen Wan, Han Zhou, Ruoxi Sun, Hootan Nakhost, Ke Jiang, Rajarishi Sinha, Sercan Ö. Arık

机构 * Google(谷歌)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures, 2 tables (22 pages, 9 figures and 3 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16320 2025-09-16 astro-ph.IM cs.LG 71%

Learning novel representations of variable sources from multi-modal $\textit{Gaia}$ data via autoencoders

P. Huijse, J. De Ridder, L. Eyer, L. Rimoldini, B. Holl, N. Chornay, J. Roquette, K. Nienartowicz, G. Jevardat de Fombelle, D. J. Fritzewski, A. Kemp, V. Vanlaer, M. Vanrespaille, H. Wang, M. I. Carnerero, C. M. Raiteri, G. Marton, M. Madarász, G. Clementini, P. Gavras, C. Aerts

机构 * Institute of Astronomy, KU Leuven, Celestijnenlaan 200D, B-3001 Leuven, Belgium Millennium Institute of Astrophysics, Nuncio Monse\ nor Sotero Sanz 100, Of. 104, Providencia, Santiago, Chile Department of Astronomy, University of Geneva, Chemin Pegasi 51, 1290 Versoix, Switzerland Department of Astronomy, University of Geneva, Chemin d’Ecogia 16, 1290 Versoix, Switzerland Sednai S\`arl, Geneva, Switzerland INAF - Osservatorio Astrofisico di Torino, Via Osservatorio 20, I-10025 Pino Torinese, Italy Konkoly Observatory, HUN-REN Research Centre for Astronomy Earth Sciences, Konkoly Thege 15-17, 1121 Budapest, Hungary CSFK, MTA Centre of Excellence, Konkoly Thege 15-17, 1121, Budapest, Hungary INAF - Osservatorio di Astrofisica e Scienza dello Spazio di Bologna, Via Piero Gobetti 93/3, Bologna 40129, Italy Starion for European Space Agency, Camino bajo del Castillo, s/n, Urbanizacion Villafranca del Castillo, Villanueva de la Ca \ n ada, 28692 Madrid, Spain Department of Astrophysics, IMAPP, Radboud University Nijmegen, PO Box 9010, 6500 GL Nijmegen, The Netherlands Max Planck Institute for Astronomy, Koenigstuhl 17, 69117 Heidelberg, Germany

专题命中 多模态生成 :multi-modal(title)

Comments Manuscript accepted on Astronomy & Astrophysics, 20 pages, 20 figures, 2 tables

Journal ref A&A 701, A150 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11698 2025-09-16 cs.CL cs.AI cs.CV cs.LG 67%

CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model

Wei-Hsin Yeh, Yu-An Su, Chih-Ning Chen, Yi-Hsueh Lin, Calvin Ku, Wen-Hsin Chiu, Min-Chun Hu, Lun-Wei Ku

机构 * Institute of Information Science, Academia Sinica(学术院信息研究所) National Tsing Hua University(国立清华大学) National Taiwan University(国立台湾大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-long.1413

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics Volume 1: Long Papers (2025) 29126-29151

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11661 2025-09-16 cs.CV cs.AI 62%

DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition

Lifei Hao, Yue Cheng, Baoqi Huang, Bing Jia, Xuandong Zhao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10845 2025-09-16 cs.CL cs.MM 62%

Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production

Liqian Feng, Lintao Wang, Kun Hu, Dehui Kong, Zhiyong Wang

机构 * School of Computer Science(计算机科学学院) School of Science(科学学院) Faculty of Information Technology(信息技术学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10345 2025-09-16 cs.CV cs.AI 62%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18985 2025-09-16 cs.LG cs.CL cs.CV 62%

STRICT: Stress Test of Rendering Images Containing Text

Tianyu Zhang, Xinyu Wang, Lu Li, Zhenghan Tai, Jijun Chi, Jingrui Tian, Hailin He, Suyuchen Wang

机构 * Mila, University of Montreal(蒙特利尔大学Mila) McGill University(麦吉尔大学) University of Pennsylvania(宾夕法尼亚大学) University of Toronto(多伦多大学) University of California, Los Angeles(加州大学洛杉矶分校) Southwestern University of Finance and Economics(西南财经大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11865 2025-09-16 cs.RO cs.AI 57%

Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer

Travis Davies, Yiqi Huang, Yunxin Liu, Xiang Chen, Huxian Liu, Luhui Hu

机构 * ZhiCheng AI(智成人工智能) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11406 2025-09-16 cs.CV 57%

No Modality Left Behind: Dynamic Model Generation for Incomplete Medical Data

Christoph Fürböck, Paul Weiser, Branko Mitic, Philipp Seeböck, Thomas Helbich, Georg Langs

机构 * Computational Imaging Research Lab(计算成像研究实验室) Department for Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Medical University of Vienna(维也纳医学大学) Comprehensive Center for Artificial Intelligence in Medicine(医学人工智能综合中心) Christian Doppler Laboratory for Machine Learning Driven Precision Imaging(机器学习驱动精准成像的克里斯蒂安·多普勒实验室) Department of Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉·A·马丁努斯生物医学成像中心) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院) Department of Radiology(放射科) Division of General and Pediatric Radiology(普通和儿童放射科部门)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted at MICCAI2025 ML-CDS Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10873 2025-09-16 cs.MM 57%

Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance

Jing Xiao, Hongfei Liu, Ruiqi Dong, Jimin Liu, Haoyong Yu

专题命中 多模态生成 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏