arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 129 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 32 篇

2504.20940 2025-10-28 physics.chem-ph cs.LG physics.comp-ph 50%

Energy-Based Coarse-Graining in Molecular Dynamics: A Flow-Based Framework without Data

Maximilian Stupp, P. S. Koutsourelakis

机构 * Professorship of Data-driven Materials Modeling School of Engineering and Design Technical University of Munich(数据驱动材料建模教授职位 工程与设计学院 慕尼黑技术大学) Munich Data Science Institute (MDSI - Core Member) Technical University of Munich(慕尼黑数据科学研究所(MDSI - 核心成员) 慕尼黑技术大学)

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21748 2025-10-28 eess.SP cs.SY eess.SY q-bio.NC 50%

Automated Tinnitus Detection Through Dual-Modality Neuroimaging: EEG Microstate Analysis and Resting-State fMRI Classification Using Deep Learning

Kiana Kiashemshaki, Sina Samieirad, Sarvenaz Erfani, Aryan Jalaeianbanayan, Nasibeh Asadi Isakan, Hossein Najafzadeh

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 11 篇

2510.12712 2025-10-28 cs.CV cs.AI 81%

Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning

Xingang Guo, Utkarsh Tyagi, Advait Gosai, Paula Vergara, Jayeon Park, Ernesto Gabriel Hernández Montoya, Chen Bo Calvin Zhang, Bin Hu, Yunzhong He, Bing Liu, Rakshith Sharma Srinivasa

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23032 2025-10-28 cs.CE 78%

P1GPT: a multi-agent LLM workflow module for multi-modal financial information analysis

Chen-Che Lu, Yun-Cheng Chou, Teng-Ruei Chen

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06278 2025-10-28 cs.RO cs.HC 78%

Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation

Tongfei Bian, Mathieu Chollet, Tanaya Guha

机构 * University of Glasgow(格拉斯哥大学) University of Glasgow School of Computer Science(格拉斯哥大学计算机科学学院)

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted by ACM Multimedia 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21809 2025-10-28 cs.CV cs.RO 70%

Embodied Navigation with Auxiliary Task of Action Description Prediction

Haru Kondoh, Asako Kanezaki

机构 * Institute of Science Tokyo(东京科学研究所) RIKEN AIP(日本科学技术研究所AIP)

专题命中 多模态Agent :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments ICCV 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21746 2025-10-28 cs.RO 67%

Avi: Action from Volumetric Inference

Harris Song, Long Le

机构 * University of California, Los Angeles(加州大学洛杉矶分校) University of pennsylvania(宾夕法尼亚大学)

专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract)

Comments NeurIPS 2025 Workshop on Embodied World Models for Decision Making. URL: https://avi-3drobot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15963 2025-10-28 cs.CV cs.AI cs.LG 62%

ESCA: Contextualizing Embodied Agents via Scene-Graph Generation

Jiani Huang, Amish Sethi, Matthew Kuo, Mayank Keoliya, Neelay Velingker, JungHo Jung, Ser-Nam Lim, Ziyang Li, Mayur Naik

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a Spotlight Paper at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15566 2025-10-28 cs.CV cs.AI 62%

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

Shaojie Zhang, Ruoceng Zhang, Pei Fu, Shaokang Wang, Jiahui Yang, Xin Du, Shiqi Cui, Bin Qin, Ying Huang, Zhenbo Luo, Jian Luan

机构 * Xiaomi Inc(小米公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21900 2025-10-28 cs.CL cs.AI 62%

Deep Literature Survey Automation with an Iterative Workflow

Hongbo Zhang, Han Cui, Yidong Wang, Yijian Tian, Qi Guo, Cunxiang Wang, Jian Wu, Chiyu Song, Yue Zhang

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院) Peking University(北京大学) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖先进研究院技术研究所)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21566 2025-10-28 cs.MA cs.CL 57%

ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem

Fangwen Wu, Zheng Wu, Jihong Wang, Yunku Chen, Ruiguang Pei, Heyuan Huang, Xin Liao, Xingyu Lou, Huarong Deng, Zhihui Fu, Weiwen Liu, Zhuosheng Zhang, Weinan Zhang, Jun Wang

机构 * Shanghai Jiao Tong University(上海交通大学) OPPO

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21669 2025-10-28 cs.AI 57%

SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

Wanxin Tian, Shijie Zhang, Kevin Zhang, Xiaowei Chi, Chunkai Fan, Junyu Lu, Yulin Luo, Qiang Zhou, Yiming Zhao, Ning Liu, Siyu Lin, Zhiyuan Qin, Xiaozhu Ju, Shanghang Zhang, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(人形机器人创新中心) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05795 2025-10-28 physics.ed-ph 50%

Creating a customisable Socratic AI physics tutor

Eugenio Tufino, Bor Gregorcic

专题命中 多模态Agent :multimodal(abstract)

Comments 7 pages, 3 figures

Journal ref Phys. Educ. 60 065037 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 15 篇

2501.17823 2025-10-28 cs.CV cs.AI cs.LG 88%

Robust Multimodal Learning via Cross-Modal Proxy Tokens

Md Kaykobad Reza, Ameya Patil, Mashhour Solh, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 28 Pages, 13 Figures, 11 Tables. Accepted by Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22829 2025-10-28 cs.CV cs.AI cs.MM 85%

LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction

Aleksandar Pramov

机构 * Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21793 2025-10-28 cs.CV cs.AI eess.IV 84%

2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection

Usman Ali, Ali Zia, Abdul Rehman, Umer Ramzan, Zohaib Hassan, Talha Sattar, Jing Wang, Wei Xiang

机构 * GIFT University(GIFT大学) La Trobe University(拉特罗布大学) Department of Primary Industries(初级产业部门)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 26th International Conference on Digital Image Computing: Techniques and Applications (DICTA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23273 2025-10-28 cs.LG cs.AI q-bio.QM 83%

A Novel Framework for Multi-Modal Protein Representation Learning

Runjie Zheng, Zhen Wang, Anjie Qiao, Jiancong Xie, Jiahua Rao, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University (SYSU)(计算机科学与工程学院,中山大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 35 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23151 2025-10-28 cs.CV cs.LG 83%

AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes

Sixian Liu, Chen Xu, Qiang Wang, Donghai Shi, Yiwen Li

机构 * Yaowu Technology Co., Ltd, Shenzhen, China(深圳优华科技有限公司,深圳,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21829 2025-10-28 cs.CV 83%

A Flow Model with Low-Rank Transformers for Incomplete Multimodal Survival Analysis

Yi Yin, Yuntao Shou, Zao Dai, Yun Peng, Tao Meng, Wei Ai, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业与技术大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22507 2025-10-28 cs.CV cs.AI 81%

GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease Diagnosis

Rui Jin, Chen Chen, Yin Liu, Hongfu Sun, Min Zeng, Min Li, Yang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally to this work. Correspondence to: Yang Gao, E-mail: yang.gao@csu.edu.cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22964 2025-10-28 cs.CV 79%

Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges

Liling Yang, Ning Chen, Jun Yue, Yidan Liu, Jiayi Ma, Pedram Ghamisi, Antonio Plaza, Leyuan Fang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Institute of Remote Sensing and Geographic Information System, Peking University(北京大学遥感与地理信息系统研究所) School of Automation, Central South University(中南大学自动化学院) Electronic Information School, Wuhan University(武汉大学电子信息学院) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍尔茨中心) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Hyperspectral Computing Laboratory, Department of Technology of Computers and Communications, Escuela Politécnica, University of Extremadura(埃斯特雷马杜拉大学技术计算机与通讯系超光谱计算实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21808 2025-10-28 cs.CV cs.AI 62%

Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning

Jiaao Yu, Mingjie Han, Jinkun Jiang, Junyu Dong, Tao Gong, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(东华大学计算机科学与技术学院) College of Computer Science and Technology, Ocean University of China, China(中国海洋大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21794 2025-10-28 cs.CV cs.AI 62%

Token-Level Inference-Time Alignment for Vision-Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Kewei Gao, Jian Lou, Zunlei Feng, Mingli Song

机构 * Zhejiang University(浙江大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 57%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22301 2025-10-28 cs.LG cs.AI 57%

AnyECG-Lab: An Exploration Study of Fine-tuning an ECG Foundation Model to Estimate Laboratory Values from Single-Lead ECG Signals

Yujie Xiao, Gongzhen Tang, Wenhui Liu, Jun Li, Guangkun Nie, Zhuoran Kan, Deyun Zhang, Qinghao Zhao, Shenda Hong

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学人民医院医学技术研究所) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) HeartVoice Medical Technology(心声医疗技术) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) State Key Laboratory of Vascular Homeostasis and Remodeling, NHC Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides, Peking University(国家心血管病分子生物学与调节肽重点实验室,北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16188 2025-10-28 cs.CV 57%

Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning

Ming Li, Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Bill Zhu, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) University of Southern California(南加州大学) Emory University(埃默里大学) Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CV

Comments Neurips 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23449 2025-10-28 cs.LG 50%

Schrodinger Neural Network and Uncertainty Quantification: Quantum Machine

M. M. Hammad

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22239 2025-10-28 eess.IV cs.LG q-bio.QM 50%

Synthetic-to-Real Transfer Learning for Chromatin-Sensitive PWS Microscopy

Jahidul Arafat, Sanjaya Poudel

机构 * Department of Computer Science and Software Engineering, Auburn University, Alabama, USA(计算机科学与软件工程系,阿伯茨罕大学,阿拉巴马州,美国)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 24 pages, 5 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 11 篇

2510.21906 2025-10-28 cs.NE cs.AI 83%

Structure-Aware Cooperative Ensemble Evolutionary Optimization on Combinatorial Problems with Multimodal Large Language Models

Jie Zhao, Kang Hao Cheong

机构 * School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11277 2025-10-28 eess.IV cs.AI cs.CV 81%

Macro2Micro: A Rapid and Precise Cross-modal Magnetic Resonance Imaging Synthesis using Multi-scale Structural Brain Similarity

Sooyoung Kim, Joonwoo Kwon, Junbeom Kwon, Jungyoun Janice Min, Sangyoon Bae, Yuewei Lin, Shinjae Yoo, Jiook Cha

机构 * Department of Brain and Cognitive Science, Seoul National University, Seoul, Republic of Korea(脑科学与认知科学系,首尔国立大学) Department of Applied Bioengineering, Seoul National University, Seoul, Republic of Korea(应用生物工程系,首尔国立大学) Department of Psychology, Seoul National University, Seoul, Republic of Korea(心理学系,首尔国立大学) Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, Republic of Korea(人工智能跨学科项目,首尔国立大学) Computational Science Initiative, Brookhaven National Laboratory, Upton, NY, USA(计算科学计划,布鲁赫斯国家实验室) School of Economics, Sogang University(经济学院,成均馆大学) Brookhaven National Laboratory(布鲁赫斯国家实验室)

专题命中 其他多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments The code will be made available upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏