arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 74 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 13 篇

2509.11082 2025-09-16 cs.CV cs.RO 79%

Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation

Zongwu Xie, Kaijie Yun, Yang Liu, Yiming Ji, Han Li

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10704 2025-09-16 cs.AI cs.CV 73%

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

Xingchen Wan, Han Zhou, Ruoxi Sun, Hootan Nakhost, Ke Jiang, Rajarishi Sinha, Sercan Ö. Arık

机构 * Google(谷歌)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures, 2 tables (22 pages, 9 figures and 3 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16320 2025-09-16 astro-ph.IM cs.LG 71%

Learning novel representations of variable sources from multi-modal $\textit{Gaia}$ data via autoencoders

P. Huijse, J. De Ridder, L. Eyer, L. Rimoldini, B. Holl, N. Chornay, J. Roquette, K. Nienartowicz, G. Jevardat de Fombelle, D. J. Fritzewski, A. Kemp, V. Vanlaer, M. Vanrespaille, H. Wang, M. I. Carnerero, C. M. Raiteri, G. Marton, M. Madarász, G. Clementini, P. Gavras, C. Aerts

机构 * Institute of Astronomy, KU Leuven, Celestijnenlaan 200D, B-3001 Leuven, Belgium Millennium Institute of Astrophysics, Nuncio Monse\ nor Sotero Sanz 100, Of. 104, Providencia, Santiago, Chile Department of Astronomy, University of Geneva, Chemin Pegasi 51, 1290 Versoix, Switzerland Department of Astronomy, University of Geneva, Chemin d’Ecogia 16, 1290 Versoix, Switzerland Sednai S\`arl, Geneva, Switzerland INAF - Osservatorio Astrofisico di Torino, Via Osservatorio 20, I-10025 Pino Torinese, Italy Konkoly Observatory, HUN-REN Research Centre for Astronomy Earth Sciences, Konkoly Thege 15-17, 1121 Budapest, Hungary CSFK, MTA Centre of Excellence, Konkoly Thege 15-17, 1121, Budapest, Hungary INAF - Osservatorio di Astrofisica e Scienza dello Spazio di Bologna, Via Piero Gobetti 93/3, Bologna 40129, Italy Starion for European Space Agency, Camino bajo del Castillo, s/n, Urbanizacion Villafranca del Castillo, Villanueva de la Ca \ n ada, 28692 Madrid, Spain Department of Astrophysics, IMAPP, Radboud University Nijmegen, PO Box 9010, 6500 GL Nijmegen, The Netherlands Max Planck Institute for Astronomy, Koenigstuhl 17, 69117 Heidelberg, Germany

专题命中 多模态生成 :multi-modal(title)

Comments Manuscript accepted on Astronomy & Astrophysics, 20 pages, 20 figures, 2 tables

Journal ref A&A 701, A150 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11698 2025-09-16 cs.CL cs.AI cs.CV cs.LG 67%

CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model

Wei-Hsin Yeh, Yu-An Su, Chih-Ning Chen, Yi-Hsueh Lin, Calvin Ku, Wen-Hsin Chiu, Min-Chun Hu, Lun-Wei Ku

机构 * Institute of Information Science, Academia Sinica(学术院信息研究所) National Tsing Hua University(国立清华大学) National Taiwan University(国立台湾大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-long.1413

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics Volume 1: Long Papers (2025) 29126-29151

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11661 2025-09-16 cs.CV cs.AI 62%

DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition

Lifei Hao, Yue Cheng, Baoqi Huang, Bing Jia, Xuandong Zhao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10845 2025-09-16 cs.CL cs.MM 62%

Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production

Liqian Feng, Lintao Wang, Kun Hu, Dehui Kong, Zhiyong Wang

机构 * School of Computer Science(计算机科学学院) School of Science(科学学院) Faculty of Information Technology(信息技术学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10345 2025-09-16 cs.CV cs.AI 62%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18985 2025-09-16 cs.LG cs.CL cs.CV 62%

STRICT: Stress Test of Rendering Images Containing Text

Tianyu Zhang, Xinyu Wang, Lu Li, Zhenghan Tai, Jijun Chi, Jingrui Tian, Hailin He, Suyuchen Wang

机构 * Mila, University of Montreal(蒙特利尔大学Mila) McGill University(麦吉尔大学) University of Pennsylvania(宾夕法尼亚大学) University of Toronto(多伦多大学) University of California, Los Angeles(加州大学洛杉矶分校) Southwestern University of Finance and Economics(西南财经大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11865 2025-09-16 cs.RO cs.AI 57%

Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer

Travis Davies, Yiqi Huang, Yunxin Liu, Xiang Chen, Huxian Liu, Luhui Hu

机构 * ZhiCheng AI(智成人工智能) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11406 2025-09-16 cs.CV 57%

No Modality Left Behind: Dynamic Model Generation for Incomplete Medical Data

Christoph Fürböck, Paul Weiser, Branko Mitic, Philipp Seeböck, Thomas Helbich, Georg Langs

机构 * Computational Imaging Research Lab(计算成像研究实验室) Department for Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Medical University of Vienna(维也纳医学大学) Comprehensive Center for Artificial Intelligence in Medicine(医学人工智能综合中心) Christian Doppler Laboratory for Machine Learning Driven Precision Imaging(机器学习驱动精准成像的克里斯蒂安·多普勒实验室) Department of Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉·A·马丁努斯生物医学成像中心) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院) Department of Radiology(放射科) Division of General and Pediatric Radiology(普通和儿童放射科部门)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted at MICCAI2025 ML-CDS Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10873 2025-09-16 cs.MM 57%

Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance

Jing Xiao, Hongfei Liu, Ruiqi Dong, Jimin Liu, Haoyong Yu

专题命中 多模态生成 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 14 篇

2509.11112 2025-09-16 cs.NI cs.AI cs.ET cs.IT cs.LG math.IT 83%

Multi-Modal Sensing Aided mmWave Beamforming for V2V Communications with Transformers

Muhammad Baqer Mollah, Honggang Wang, Hua Fang

机构 * University of Massachusetts Dartmouth(马萨诸塞大学达特茅斯分校) Yeshiva University(耶鲁大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 6 Pages, Accepted to present at 2025 IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06516 2025-09-16 cs.LG cs.AI 83%

QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients

Zongheng Guo, Tao Chen, Manuela Ferrario

机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程系,米兰理工学院) State Key Laboratory of Industrial Control Technology, Zhejiang University(工业控制技术国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

Comments 11 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07030 2025-09-16 cs.CL cs.AI cs.CV cs.IR cs.LG 82%

FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering

Amirhossein Abaskohi, Spandana Gella, Giuseppe Carenini, Issam H. Laradji

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学) ServiceNow Research(ServiceNow研究)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11620 2025-09-16 cs.CL cs.CY 79%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02006 2025-09-16 cs.AI 79%

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang, Yunchao Wei, Meng Fang, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) The Pennsylvania State University(宾夕法尼亚州立大学) Beijing Jiaotong University(北京交通大学) University of Liverpool(利物浦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11335 2025-09-16 cs.LG cond-mat.mtrl-sci 78%

MatQnA: A Benchmark Dataset for Multi-modal Large Language Models in Materials Characterization and Analysis

Yonghao Weng, Liqiang Gao, Linwu Zhu, Jian Huang

机构 * Department of Materials Engineering(材料工程系) Zhejiang University(浙江大学) Department of Data Intelligence(数据智能系) Shiyanjia Lab of Scientific Compass(科学之桥实验室)

专题命中 多模态评测 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20358 2025-09-16 cs.LG 78%

Developing a Multi-Modal Machine Learning Model For Predicting Performance of Automotive Hood Frames

Abhishek Indupally, Satchit Ramnath

专题命中 多模态评测 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11136 2025-09-16 cs.LG cs.AI cs.CL cs.SI 76%

Agentic Username Suggestion and Multimodal Gender Detection in Online Platforms: Introducing the PNGT-26K Dataset

Farbod Bijary, Mohsen Ebadpour, Amirhosein Tajbakhsh

机构 * Amirkabir University of Technology(阿米尔卡比尔理工大学) Iran University of Science & Technology(伊朗科学技术大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10683 2025-09-16 cs.CV cs.AI 62%

A Comparison and Evaluation of Fine-tuned Convolutional Neural Networks to Large Language Models for Image Classification and Segmentation of Brain Tumors on MRI

Felicia Liu, Jay J. Yoo, Farzad Khalvati

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 57%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11459 2025-09-16 cs.AI 57%

Knowledge-Guided Adaptive Mixture of Experts for Precipitation Prediction

Chen Jiang, Kofi Osei, Sai Deepthi Yeddula, Dongji Feng, Wei-Shinn Ku

机构 * Samuel Ginn College of Engineering(萨姆uel吉恩工程学院) Auburn University(阿伯杜大学) Gustavus Adolphus College(古斯塔夫·阿道夫学院) School of Engineering and Computer Science(工程与计算机科学学院) Oakland University(奥克兰大学) School of Computing and Design(计算与设计学院) California State University Monterey Bay(蒙特利尔湾加州州立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10570 2025-09-16 cs.RO cs.AI 57%

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang, Sisuo Lyu, Haicheng Liao, Limin Yu, Weiping Ding, Runwei Guan, Yutao Yue

机构 * Department of Mathematical Sciences, School of Physical sciences, University of Liverpool(利物浦大学数学科学系) Department of Communications and Networking, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学通讯与网络系) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Deep Interdisciplinary Intelligence Lab, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)深度跨学科智能实验室) School of Mathematics and Statistics, Shandong University(山东大学数学与统计学院) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)数据科学与分析方向) Institute of Deep Perception Technology, Jiangsu(江苏深度感知技术研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11876 2025-09-16 cs.HC 50%

Lost in Data: How Older Adults Perceive and Navigate Health Data Representations

Peterson Jean, Emma Murphy, Enda Bates

专题命中 多模态评测 :multimodal(abstract)

Comments AAATE 2025 Proceedings (Research Strand). Licensed under CC BY-NC-ND 4.0. ISBN: 978-9925-604-07-4

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10556 2025-09-16 q-bio.TO cs.CE 50%

COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis

Nina Wiedemann, Dianne de Korte-de Boer, Matthias Richter, Sjors van de Weijer, Charlotte Buhre, Franz A. M. Eggert, Sophie Aarnoudse, Lotte Grevendonk, Steffen Röber, Carlijn M. E. Remie, Wolfgang Buhre, Ronald Henry, Jannis Born

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 6 篇

2509.11270 2025-09-16 cs.RO cs.AI 79%

Embodied Intelligence in Disassembly: Multimodal Perception Cross-validation and Continual Learning in Neuro-Symbolic TAMP

Ziwen He, Zhigang Wang, Yanlong Peng, Pengxu Chang, Hong Yang, Ming Chen

机构 * School of Mechanical Engineering, Shanghai Jiao Tong University(上海交通大学机械工程学院) Intel Labs China(英特尔中国实验室) Intel Asia Pacific R&D Ltd(英特尔亚太研发有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 3 figures. Accepted at CASE2025. This arXiv version contains minor corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11782 2025-09-16 cs.LG q-bio.BM 78%

Multimodal Regression for Enzyme Turnover Rates Prediction

Bozhen Hu, Cheng Tan, Siyuan Li, Jiangbin Zheng, Sizhe Qiu, Jun Xia, Stan Z. Li

机构 * AI Division, School of Engineering, Westlake University(西拉丘学院人工智能系,西湖大学) Zhejiang University(浙江大学) Oxford University(牛津大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 9 pages, 5 figures. This paper was withdrawn from the IJCAI 2025 proceedings due to the lack of participation in the conference and presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10576 2025-09-16 cs.CY cs.AI 57%

Aesthetic Experience and Educational Value in Co-creating Art with Generative AI: Evidence from a Survey of Young Learners

Chengyuan Zhang, Suzhe Xu

机构 * Huaqiao University(华侨大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11793 2025-09-16 cs.RO 50%

UniPilot: Enabling GPS-Denied Autonomy Across Embodiments

Mihir Kulkarni, Mihir Dharmadhikari, Nikhil Khedekar, Morten Nissov, Mohit Singh, Philipp Weiss, Kostas Alexis

机构 * Department of Engineering Cybernetics, O. S. Bragstads Plass 2D, Norwegian University of Science and Technology (NTNU), Trondheim, Norway(工程 cybernetics 系,挪威科学技术大学(NTNU),特隆赫姆,挪威)

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11516 2025-09-16 cs.RO cs.SY eess.SY 50%

PaiP: An Operational Aware Interactive Planner for Unknown Cabinet Environments

Chengjin Wang, Zheng Yan, Yanmin Zhou, Runjie Shen, Zhipeng Wang, Bin Cheng, Bin He

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏