arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-20 至 2025-11-20 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 10 篇

2411.14886 2025-11-20 eess.SP cs.LG 78%

Abnormality Prediction and Forecasting of Laboratory Values from Electrocardiogram Signals Using Multimodal Deep Learning

Juan Miguel Lopez Alcaraz, Nils Strodthoff

机构 * Carl von Ossietzky Universität Oldenburg(奥尔登堡卡尔·冯·奥西特齐克大学) AI4Health Division(AI4Health部门)

专题命中 多模态评测 :multimodal(title,abstract)

Comments Accepted for publication in Scientific Reports. 15 pages, 2 figures. Code available at: https://github.com/AI4HealthUOL/CardioLab

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13869 2025-11-20 cs.CV cs.AI 62%

H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction

Xueyang Li, Zongren Wang, Yuliang Zhang, Zixuan Pan, Yu-Jen Chen, Nishchal Sapkota, Gelei Xu, Danny Z. Chen, Yiyu Shi

机构 * Engineering, University of Notre Dame, USA(工程学院,诺特大学,美国) Department of Urology, The First Affiliated Hospital, Sun Yat-sen University, China(泌尿科,中山大学附属第一医院,中国)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15183 2025-11-20 cs.CL cs.LG 57%

HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples

Rishikant Chigrupaatii, Ponnada Sai Tulasi Kanishka, Lalit Chandra Routhu, Martin Patel Sama Supratheek Reddy, Divyam Gupta, Dasari Srikar, Krishna Teja Kuchimanchi, Rajiv Misra, Rohun Tripathi

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15179 2025-11-20 cs.CV 57%

MMCM: Multimodality-aware Metric using Clustering-based Modes for Probabilistic Human Motion Prediction

Kyotaro Tokoro, Hiromu Taketsugu, Norimichi Ukita

机构 * Toyota Technological Institute(丰田技术研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13326 2025-11-20 stat.AP cs.AI 57%

TacEleven: generative tactic discovery for football open play

Siyao Zhao, Hao Ma, Zhiqiang Pu, Jingjing Huang, Yi Pan, Shijie Wang, Zhi Ming

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(交叉科学学院,中国科学院大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shanghai AI Laboratory(上海人工智能实验室) Association de la Jeunesse Auxerroise(亚眠青年协会)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2511.15370 2025-11-20 cs.CL cs.AI 62%

The Empowerment of Science of Science by Large Language Models: New Tools and Methods

Guoqiang Liang, Jingqian Gong, Mengxuan Li, Gege Lin, Shuo Zhang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments The manuscript is currently ongoing the underreview process of the journal of information science

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15456 2025-11-20 cs.AI q-fin.GN 57%

Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining

了解您的意图:一种自主多视角大语言模型代理框架用于DeFi用户交易意图挖掘

Qian'ang Mao, Yuxuan Zhang, Jiaman Chen, Wenjun Zhou, Jiaqi Yan

机构 * Nanjing University(南京大学) The University of Tennessee, Knoxville(田纳西大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出TIM框架,通过多视角LLM代理系统挖掘DeFi用户交易意图,提升意图推断的准确性和可验证性。

Comments Written in 2025 Q1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23596 2025-11-20 cs.AI 57%

Agent-SAMA: State-Aware Mobile Assistant

Linqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun, Chen, Yang Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to AAAI-26 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 16 篇

2511.14766 2025-11-20 cs.IR cs.MM 83%

OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction

Yang Li, Yajiao Wang, Wenhao Hu, Zhixiong Zhang, Mengting Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14969 2025-11-20 eess.AS cs.AI cs.LG eess.IV eess.SP 81%

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

Zanxu Wang, Homayoon Beigi

机构 * Columbia University, New York, USA(哥伦比亚大学) Recognition Technologies, Inc., New York, USA(识别技术公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments 8 pages, 14 images, 3 tables, Recognition Technologies, Inc. Technical Report RTI-20251118-01

Journal ref Recognition Technologies, Inc. Technical Reports, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15433 2025-11-20 cs.CV 79%

Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection

基于模态解耦的表示空间约束学习用于多模态目标检测

YiKang Shao, Tao Shi

机构 * school of reliability and systems engineering, Beihang University(可靠性与系统工程学院,北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出RSC-MD方法,通过模态解耦和表示空间约束学习解决多模态目标检测中的融合退化问题,提升各模态的优化效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12079 2025-11-20 cs.CV 79%

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

Hongxuan Li, Wencheng Zhu, Huiying Xu, Xinzhong Zhu, Pengfei Zhu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2026. 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04638 2025-11-20 cs.CV 79%

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

Xixi Wan, Aihua Zheng, Bo Jiang, Beibei Wang, Chenglong Li, Jin Tang

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province, School of Artificial Intelligence, Anhui University(安徽省信息材料与智能感知实验室,人工智能学院,安徽大学) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University(安徽省多模态认知计算重点实验室,计算机科学与技术学院,安徽大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15138 2025-11-20 cs.LG cs.HC 78%

Cross-Modal Consistency-Guided Active Learning for Affective BCI Systems

Hyo-Jeong Jang, Hye-Bin Shin, Kang Yin

机构 * Dept. of Brain and Cognitive Engineering(脑科学与认知工程系) Korea University(韩国大学) Dept. of Artificial Intelligence(人工智能系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15139 2025-11-20 q-bio.GN cs.AI cs.LG 74%

CASPER: Cross-modal Alignment of Spatial and single-cell Profiles for Expression Recovery

Amit Kumar, Maninder Kaur, Raghvendra Mall, Sukrit Gupta

机构 * Department of Computer Science \& Engineering, Indian Institute of Technology Ropar, India Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar. Department of Biomedical Engineering, Indian Institute of Technology Ropar, India

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14970 2025-11-20 cs.CV cs.AI cs.RO 62%

EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects

Gbenga Omotara, Ramy Farag, Seyed Mohamad Ali Tousi, G. N. DeSouza

机构 * Vision-Guided and Intelligent Robotics Lab (ViGIR) University of Missouri(视觉引导与智能机器人实验室(ViGIR)大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00275 2025-11-20 cs.CV cs.AI 62%

AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding

Md Asaduzzaman Jabin, Hanqi Jiang, Yiwei Li, Patrick Kaggwa, Eugene Douglass, Juliet N. Sekandi, Tianming Liu

机构 * University of Georgia(佐治亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 7th International Workshop on Large Scale Holistic Video Understanding: Toward Video Foundation Models

Journal ref Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07033 2025-11-20 cs.CV 57%

One Latent Space to Rule All Degradations: Unifying Restoration Knowledge for Image Fusion

Haolong Ma, Hui Li, Chunyang Cheng, Zeyang Zhang, Xiaoqing Luo, Xiaoning Song, Xiao-Jun Wu

机构 * Jiangnan University(江南大学) Suzhou University of Science and Technology(苏州科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15016 2025-11-20 cs.CV 57%

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

Zhenyu Cui, Jiahuan Zhou, Yuxin Peng

机构 * Zhenyu Cui, Jiahuan Zhou, Yuxin Peng(作者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05923 2025-11-20 cs.CV 57%

Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 57%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18757 2025-11-20 cs.CV 57%

ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * CCDS, NTU, Singapore(南洋理工大学新加坡分校) CCST, ZJUT, China(浙江工业大学计算机科学与技术学院) Terminus AI Lab, UCAS, China(中国科学院大学人工智能实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15251 2025-11-20 cs.LG cs.NI 50%

PLATONT: Learning a Platonic Representation for Unified Network Tomography

Chengze Du, Heng Xu, Zhiwei Yu, Bo Liu, Jialong Li

机构 * Computer Science and Control Engineering, Shenzhen University of Advanced Technology(深圳先进技术大学计算机科学与控制工程系) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14988 2025-11-20 cs.RO 50%

An Alignment-Based Approach to Learning Motions from Demonstrations

Alex Cuellar, Christopher K Fourie, Julie A Shah

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 8 pages, 8 figures, originally published in the IEEE Robotics and Automation Letters

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 11, pp. 11912-11919, Nov. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 6 篇

2511.15600 2025-11-20 cs.CV cs.LG 79%

US-X Complete: A Multi-Modal Approach to Anatomical 3D Shape Recovery

US-X Complete: 一种多模态方法用于解剖三维形状恢复

Miruna-Alexandra Gafencu, Yordanka Velikova, Nassir Navab, Mohammad Farid Azampour

机构 * Computer-Aided Medical Procedures (CAMP), Technical University of Munich, Munich, Germany(计算机辅助医学程序(CAMP),慕尼黑技术大学,慕尼黑,德国) Munich Center for Machine Learning (MCML), Germany(慕尼黑机器学习中心(MCML),德国) Konrad Zuse School of Excellence in Reliable AI (relAI), Germany(康拉德·祖斯卓越可靠人工智能学校(relAI),德国)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 US-X Complete通过结合X光图像与超声数据,提升3D超声中椎体结构的重建精度,实现更完整的椎体可视化。

Comments Accepted at the Workshop on Shape in Medical Imaging at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15375 2025-11-20 cs.LG cs.AI 57%

Parameter Importance-Driven Continual Learning for Foundation Models

Lingxiang Wang, Hainan Zhang, Zhiming Zheng

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15194 2025-11-20 cs.RO cs.AI 57%

Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization

Jian Deng, Yuandong Wang, Yangfu Zhu, Tao Feng, Tianyu Wo, Zhenzhou Shao

机构 * Beijing Key Laboratory of Light Industrial Robot and Safety Verification, Capital Normal University(北京工业机器人与安全验证重点实验室,首都师范大学) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Software, Beihang University(北航软件学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 12 pages, 4 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12001 2025-11-20 cs.CL cs.HC 57%

Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, Motahhare Eslami

机构 * Seoul National University(首尔国立大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Under review; 16 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04599 2025-11-20 stat.ME math.ST stat.ML stat.TH 50%

From Global to Local Correlation: Geometric Decomposition of Statistical Inference

Pawel Gajer, Jacques Ravel

专题命中 其他多模态 :multi-modal(abstract)

Comments 58 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15196 2025-11-20 stat.ML cs.LG hep-lat 50%

Particle Monte Carlo methods for Lattice Field Theory

David Yallup

机构 * Kavli Institute for Cosmology Cambridge, University of Cambridge(卡弗里宇宙学研究所剑桥,剑桥大学)

专题命中 其他多模态 :multimodal(abstract)

Comments To appear in the NeurIPS 2025 workshop, Frontiers in Probabilistic Inference: Sampling Meets Learning

详情

展开后加载摘要…

URL PDF HTML 收藏