arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-17 至 2025-10-17 共收录 54 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2510.14307 2025-10-17 cs.CL cs.AI 81%

MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) Defence Science and Technology Agency(国防科学与技术局)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11817 2025-10-17 cs.CV 79%

GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models

Jonathan Roberts, Kai Han, Samuel Albanie

机构 * University of Cambridge(剑桥大学) The University of Hong Kong(香港大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21851 2025-10-17 cs.CV 79%

On Large Multimodal Models as Open-World Image Classifiers

Alessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang, Paolo Rota, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV2025, 23 pages, 13 figures, code is available at https://github.com/altndrr/lmms-owc

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14905 2025-10-17 cs.CV 79%

Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation

Siwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang, Zichen Wen, Yize Chen, Jiang Wu, Wenjun Wu, Conghui He, Weijia Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sun Yat-Sen University(中山大学) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(北京未来区块链与隐私计算高级创新中心,北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17040 2025-10-17 cs.CV cs.AI 73%

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

Hang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng, Zhenyan Chen, Chenyang Zhang, Wang Xiao, Shan Huang, Yuqi Pan, Tao Qi, Sicong Leng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14621 2025-10-17 cs.AI cs.CL 62%

ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks

Yuanyi Song, Heyuan Huang, Qiqiang Lin, Yin Zhao, Xiangmou Qu, Jun Wang, Xingyu Lou, Weiwen Liu, Zhuosheng Zhang, Jun Wang, Yong Yu, Weinan Zhang, Zhaoxiang Wang

机构 * Shanghai Jiao Tong University(上海交通大学) OPPO

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14443 2025-10-17 cs.SD cs.AI eess.AS 62%

Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare

Mayuri Kate, Suresh Neethirajan

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 40 pages, 14 figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14532 2025-10-17 cs.CV 57%

Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology

Xinrui Huang, Fan Xiao, Dongming He, Anqi Gao, Dandan Li, Xiaofan Zhang, Shaoting Zhang, Xudong Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine, Department of Oral Craniomaxillofacial(上海第九人民医院,上海交通大学医学院,口腔颅面医学部) Shanghai Jiao Tong University, School of Computer Science(上海交通大学,计算机科学学院) University of Electronic Science and Technology of China, School of Mechanical and Electrical Engineering(电子科技大学,机械与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University, College of Stomatology(上海交通大学,口腔医学院) National Center for Stomatology(国家口腔医学中心) National Clinical Medical Research Center for Oral Diseases(国家口腔疾病临床医学研究中心) Shanghai Key Laboratory of Stomatology(上海口腔医学重点实验室) Shanghai Research Institute of Stomatology(上海口腔研究所) Chinese Academy of Medical Science, Research Unit of Oral and Maxillofacial Regenerative Medicine(中国医学科学院,口腔颌面再生医学研究单元)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02535 2025-10-17 cs.CY cs.AI 57%

PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations

Rifaa Qadri, Anh Nhat Nhu, Swati Ramnath, Laura Yu Zheng, Raj Bhansali, Sylvette La Touche-Howard, Tracy Marie Zeeger, Tom Goldstein, Ming Lin

机构 * University of Maryland(马里兰大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2510.13845 2025-10-17 q-bio.NC 82%

Embodiment in multimodal large language models

Akila Kadambi, Lisa Aziz-Zadeh, Antonio Damasio, Marco Iacoboni, Srini Narayanan

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19905 2025-10-17 cs.AI 79%

EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM

Shuang Ao, Flora D. Salim, Simon Khan

机构 * University of New South Wales(新南威尔士大学) Air Force Research Laboratory(空军研究实验室)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14194 2025-10-17 cs.AI 57%

Implementation of AI in Precision Medicine

Göktuğ Bender, Samer Faraj, Anand Bhardwaj

机构 * Desautels Faculty of Management(德萨尔斯管理学院) McGill University(麦吉尔大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04241 2025-10-17 cs.HC 50%

RunPacer: A Smartwatch-Based Vibrotactile Feedback System for Symmetric Co-Running by Visually Impaired Individuals and Guides

Yichen Yu, Huan-Song Xu, Ming-Yen Lin

专题命中 多模态Agent :multimodal(abstract)

Comments 6 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 7 篇

2403.15356 2025-10-17 cs.CV 86%

Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation

Zhitong Xiong, Yi Wang, Fahong Zhang, Adam J. Stewart, Joëlle Hanna, Damian Borth, Ioannis Papoutsis, Bertrand Le Saux, Gustau Camps-Valls, Xiao Xiang Zhu

机构 * Chair of Data Science in Earth Observation, Technical University of Munich (TUM)(地球观测数据科学教授职位,慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) AIML Lab, School of Computer Science, University of St. Gallen(人工智能实验室,圣加尔登大学计算机科学学院) School of Rural, Surveying and Geoinformatics Engineering, National Technical University of Athens(农村、测绘与地理信息工程学院,国家技术大学雅典) Image Processing Laboratory (IPL), Universitat de València(图像处理实验室(IPL),瓦伦西亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14344 2025-10-17 cs.CR cs.AI 79%

BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection

Zichen Liu, Shao Yang, Xusheng Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20401 2025-10-17 cs.CV cs.RO 79%

SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment

Binod Singh, Sayan Deb Sarkar, Iro Armeni

机构 * Technical University of Munich(慕尼黑技术大学) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Project Page: https://singhbino3d.github.io/sgpp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14411 2025-10-17 cs.LG cs.MM cs.SD eess.AS 74%

Revisit Modality Imbalance at the Decision Layer

Xiaoyu Ma, Hao Chen

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),教育部,中国)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Some Insights in Balanced Multimodal Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14387 2025-10-17 cs.AI 70%

Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?

Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang, Qiufeng Wang

机构 * Xi’an-Jiaotong Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学) Duke Kunshan University(杜克昆山大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14374 2025-10-17 cs.CV 70%

Spatial Preference Rewarding for MLLMs Spatial Understanding

Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory(上海人工智能实验室) Sensetime Research(商汤科技研究院) Zhejiang University of Technology(浙江工业大学) UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01411 2025-10-17 cs.CV cs.AI 62%

ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition

Minjeong Park, Hongbeen Park, Jinkyu Kim

机构 * Department of Computer Science and Engineering, Korea University, Seoul 02841, Korea(计算机科学与工程系,韩国大学,首尔02841,韩国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 4 篇

2510.14340 2025-10-17 eess.IV cs.AI cs.CV cs.LG 81%

A Density-Informed Multimodal Artificial Intelligence Framework for Improving Breast Cancer Detection Across All Breast Densities

Siva Teja Kakileti, Bharath Govindaraju, Sudhakar Sampangi, Geetha Manjunath

专题命中 其他多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13660 2025-10-17 cs.CV 57%

OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild

Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) Zhejiang University(浙江大学) Nanjing Forestry University(南京林业大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025; Project page: https://github.com/quhongyu/OmniGaze

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00527 2025-10-17 cs.RO cs.AI 57%

Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions

Guanwen Xie, Jingzehua Xu, Yimian Ding, Zhi Zhang, Shuai Zhang, Yi Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, 518055, China(清华大学深圳国际研究生院,清华大学,深圳,518055,中国) Department of Data Science, New Jersey Institute of Technology(数据科学系,新泽西理工学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted by IEEE/RSJ IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14445 2025-10-17 cs.LG physics.geo-ph 50%

Towards geological inference with process-based and deep generative modeling, part 1: training on fluvial deposits

Guillaume Rongier, Luk Peeters

专题命中 其他多模态 :multimodal(abstract)

Comments 24 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏