arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 110 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 25 篇

2510.19954 2025-11-04 cs.AI cs.DB cs.LG 79%

RELATE: A Schema-Agnostic Perceiver Encoder for Multimodal Relational Graphs

Joe Meyer, Divyansha Lachi, Mahmoud Mohammadi, Roshan Reddy Upendra, Eva L. Dyer, Mark Li, Tom Palczewski

机构 * SAP Palo Alto, CA, USA(SAP帕洛阿尔托分校) University of Pennsylvania(宾夕法尼亚大学) SAP Seattle, WA, USA(SAP西雅图分校)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01186 2025-11-04 cs.RO cs.CV 79%

LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping

Lijie Wang, Lianjie Guo, Ziyi Xu, Qianhao Wang, Fei Gao, Xieyuanli Chen

机构 * State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University(工业控制技术国家重点实验室,系统与控制研究院,浙江大学) Differential Robot Technology Co., Ltd.(差分机器人技术有限公司) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00584 2025-11-04 cs.IR cs.CL 79%

Structurally Refined Graph Transformer for Multimodal Recommendation

Ke Shi, Yan Zhang, Miao Zhang, Lifan Chen, Jiali Yi, Kui Xiao, Xiaoju Hou, Zhifei Li

机构 * School of Computer Science, Hubei University(湖北大学计算机学院) Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University(湖北大学大数据智能分析与应用重点实验室) Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education(智能传感系统与安全重点实验室(湖北大学)) Institute of Vocational Education, Guangdong Industry Polytechnic University(广东职业技术学院教育学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Comment: 13 pages, 7 figures, accepted by IEEE Transactions on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00424 2025-11-04 cs.AI 79%

A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method

Ashutosh Anshul, Gumpili Sai Pranav, Mohammad Zia Ur Rehman, Nagendra Kumar

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Journal ref IEEE Transactions on Computational Social Systems, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00389 2025-11-04 cs.CV 79%

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

Fan Zhang, Haoxuan Li, Shengju Qian, Xin Wang, Zheng Lian, Hao Wu, Zhihong Zhu, Yuan Gao, Qiankun Li, Yefeng Zheng, Zhouchen Lin, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学) Tencent(腾讯) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01406 2025-11-04 eess.SP 78%

AoI-Aware Machine Learning for Constrained Multimodal Sensing-Aided Communications

Abolfazl Zakeri, Nhan Thanh Nguyen, Ahmed Alkhateeb, Markku Juntti

专题命中 多模态评测 :multimodal(title,abstract)

Comments Submitted to an IEEE conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01318 2025-11-04 cs.CE 78%

CSMD: Curated Multimodal Dataset for Chinese Stock Analysis

Yu Liu, Zhuoying Li, Ruifeng Yang, Fengran Mo, Cen Chen

专题命中 多模态评测 :multimodal(title,abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22810 2025-11-04 cs.CV 70%

VidText: Towards Comprehensive Evaluation for Video Text Understanding

Zhoufaran Yang, Yan Shu, Jing Wang, Zhifei Yang, Yan Zhang, Yu Li, Keyang Lu, Gangyan Zeng, Shaohui Liu, Yu Zhou, Nicu Sebe

机构 * UNITN HIT(哈尔滨工业大学) SEU(上海电子大学) PKU(北京大学) IIE, CAS(中国科学院信息工程研究所) UCAS(中国科学院大学) BUAA(北京航空航天大学) NJUST(南京理工大学) NKU(宁夏大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08221 2025-11-04 cs.CV cs.AI cs.MM 67%

EgoBlind: Towards Egocentric Visual Assistance for the Blind

Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Communication University of China(中国传媒大学) University of Science and Technology of China(中国科学技术大学) Hefei University of Technoloy(合肥工业大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments NeurIPS'25 (D&B Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17050 2025-11-04 cs.CL cs.AI cs.CE cs.CY cs.MM 67%

Towards Robust Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning

Xinyi Wu, Yanhao Jia, Qinglin Zhang, Yiran Qin, Luwei Xiao, Shuai Zhao

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00509 2025-11-04 cs.AI cs.CL cs.LG 62%

Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?

Kai Yan, Yufei Xu, Zhengyin Du, Xuesong Yao, Zheyu Wang, Xiaowen Guo, Jiecao Chen

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 3 figures, 13 tables. The paper is accepted at AACL-IJCNLP 2025 (main track), and the latest version adds modifications in camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01000 2025-11-04 cs.CV cs.LG 57%

Integrating Visual and X-Ray Machine Learning Features in the Study of Paintings by Goya

Hassan Ugail, Ismail Lujain Jaleel

机构 * Centre for Visual Computing and Intelligent Systems(视觉计算与智能系统中心) University of Bradford(布拉德福德大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00598 2025-11-04 eess.IV cs.CV 57%

GDROS: A Geometry-Guided Dense Registration Framework for Optical-SAR Images under Large Geometric Transformations

Zixuan Sun, Shuaifeng Zhi, Ruize Li, Jingyuan Xia, Yongxiang Liu, Weidong Jiang

机构 * College of Electronic Science and Technology, National University of Defense Technology(电子科学与技术学院,国防科技大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

Comments To be published in IEEE Transactions on Geoscience and Remote Sensing (T-GRS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25173 2025-11-04 cs.CV 57%

D$^2$GS: Dense Depth Regularization for LiDAR-free Urban Scene Reconstruction

Kejing Xia, Jidong Jia, Ke Jin, Yucai Bai, Li Sun, Dacheng Tao, Youjian Zhang

机构 * Wuhan University(武汉大学) Shanghai Jiaotong University(上海交通大学) TongJi University(同济大学) Bosch(博世) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10797 2025-11-04 physics.med-ph cs.CV 57%

Modality-AGnostic Image Cascade (MAGIC) for Multi-Modality Cardiac Substructure Segmentation

Nicholas Summerfield, Qisheng He, Alex Kuo, Ahmed I. Ghanem, Simeng Zhu, Chase Ruff, Joshua Pan, Anudeep Kumar, Prashant Nagpal, Jiwei Zhao, Ming Dong, Carri K. Glide-Hurst

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07769 2025-11-04 cs.CV 57%

BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled Aldahmani, Fahad Khan, Rao Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed大学人工智能学院) Linköping University(林肯大学) Shaikh Tahnoon bin Mohammed Medical City(Shaikh Tahnoon bin Mohammed医疗城) Tawam Hospital(Tawam医院) Sheikh Shakhbout Medical City(Sheikh Shakhbout医疗城) Govt Medical College Kozhikode(科钦政府医学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (Findings)

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 14051-14071

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04944 2025-11-04 cs.RO cs.LG 50%

MarsLGPR: Mars Rover Localization with Ground Penetrating Radar

Anja Sheppard, Katherine A. Skinner

机构 * University of Michigan Robotics Department(密歇根大学机器人学系) University of Michigan Space Institute(密歇根大学太空研究所)

专题命中 多模态评测 :multi-modal(abstract)

Comments IEEE Transactions on Field Robotics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06289 2025-11-04 q-fin.PM cs.LG q-fin.PR 50%

Automate Strategy Finding with LLM in Quant Investment

Zhizhuo Kou, Holam Yu, Junyu Luo, Jingshu Peng, Xujia Li, Chengzhong Liu, Juntao Dai, Lei Chen, Sirui Han, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Peking University(北京大学)

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00042 2025-11-04 cs.NE 50%

Bio-Inspired Neuron Synapse Optimization for Adaptive Learning and Smart Decision-Making

Sreeja Singh, Tamal Ghosh

专题命中 多模态评测 :multimodal(abstract)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 6 篇

2511.00940 2025-11-04 cs.RO cs.AI 83%

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

Zhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu, Che Xu, Ying Li, Chengkai Hou, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) University of Washington(华盛顿大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05335 2025-11-04 cs.CV 74%

New multimodal similarity measure for image registration via modeling local functional dependence with linear combination of learned basis functions

Joel Honkamaa, Pekka Marttinen

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Improved experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15685 2025-11-04 cs.RO 67%

From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems

Xiuchao Sui, Daiying Tian, Qi Sun, Ruirui Chen, Dongkyu Choi, Kenneth Kwok, Soujanya Poria

机构 * IHPC, Agency for Science, Technology and Research, Singapore(科技研究局智能技术中心,新加坡) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract)

Comments EMNLP 2025 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00096 2025-11-04 cs.MA cs.AI cs.CY 57%

Urban-MAS: Human-Centered Urban Prediction with LLM-Based Multi-Agent System

Shangyu Lou

机构 * University of California, Santa Barbara \& San Diego State University California USA University of California, Santa Barbara \& San Diego State University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to The 3rd ACM SIGSPATIAL International Workshop on Advances in Urban AI (UrbanAI'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00933 2025-11-04 cs.RO cs.CV 57%

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning, the University of Adelaide(澳大利亚机器学习研究所、阿德莱德大学) CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院CREATE实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00936 2025-11-04 cs.HC 50%

Exploring Human-AI Interaction with Patient-Generated Health Data Sensemaking for Cardiac Risk Reduction

Pavithren V S Pakianathan, Rania Islambouli, Hannah McGowan, Diogo Branco, Tiago Guerreiro, Jan David Smeddinck

专题命中 多模态Agent :multi-modal(abstract)

Comments Presented as demonstration at the workshop on visual analytics in healthcare (VAHC) (in conjunction with IEEE VIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 16 篇

2511.01444 2025-11-04 cs.AI 83%

Robust Multimodal Sentiment Analysis via Double Information Bottleneck

Huiting Huang, Tieliang Gong, Kai He, Jialun Wu, Erik Cambria, Mengling Feng

机构 * School of Computer Science(计算机科学学院) Technology, Xi’an Jiaotong University(技术,西安交通大学) Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(大数据知识工程省级重点实验室,西安交通大学) Saw Swee Hock School of Public Health, National University of Singapore(Saw Swee Hock 公共卫生学院,新加坡国立大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) School of Computer Science, Northwestern Polytechnical University(计算机科学学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01320 2025-11-04 cs.AI 83%

OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance

Ziqi Wang, Hailiang Zhao, Yuhao Yang, Daojiang Hu, Cheng Bao, Mingyi Liu, Kai Di, Schahram Dustdar, Zhongjie Wang, Shuiguang Deng

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Hangzhou School of Automation, Zhejiang Normal University(浙江师范大学杭州自动化学院) Distributed Systems Group at the TU Wien and with ICREA at the UPF, Barcelona(维也纳大学分布式系统组和巴塞罗那大学ICREA)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 82%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00987 2025-11-04 cs.LG 82%

Balanced Multimodal Learning via Mutual Information

Rongrong Xie, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究学院(SISSA))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01435 2025-11-04 cs.CV 79%

Contrast-Guided Cross-Modal Distillation for Thermal Object Detection

SiWoo Kim, JhongHyun An

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏