arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-20 至 2025-11-20 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2511.14766 2025-11-20 cs.IR cs.MM 83%

OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction

Yang Li, Yajiao Wang, Wenhao Hu, Zhixiong Zhang, Mengting Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14969 2025-11-20 eess.AS cs.AI cs.LG eess.IV eess.SP 81%

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

Zanxu Wang, Homayoon Beigi

机构 * Columbia University, New York, USA(哥伦比亚大学) Recognition Technologies, Inc., New York, USA(识别技术公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments 8 pages, 14 images, 3 tables, Recognition Technologies, Inc. Technical Report RTI-20251118-01

Journal ref Recognition Technologies, Inc. Technical Reports, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15433 2025-11-20 cs.CV 79%

Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection

基于模态解耦的表示空间约束学习用于多模态目标检测

YiKang Shao, Tao Shi

机构 * school of reliability and systems engineering, Beihang University(可靠性与系统工程学院,北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出RSC-MD方法,通过模态解耦和表示空间约束学习解决多模态目标检测中的融合退化问题,提升各模态的优化效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12079 2025-11-20 cs.CV 79%

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

Hongxuan Li, Wencheng Zhu, Huiying Xu, Xinzhong Zhu, Pengfei Zhu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2026. 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04638 2025-11-20 cs.CV 79%

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

Xixi Wan, Aihua Zheng, Bo Jiang, Beibei Wang, Chenglong Li, Jin Tang

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province, School of Artificial Intelligence, Anhui University(安徽省信息材料与智能感知实验室,人工智能学院,安徽大学) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University(安徽省多模态认知计算重点实验室,计算机科学与技术学院,安徽大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15138 2025-11-20 cs.LG cs.HC 78%

Cross-Modal Consistency-Guided Active Learning for Affective BCI Systems

Hyo-Jeong Jang, Hye-Bin Shin, Kang Yin

机构 * Dept. of Brain and Cognitive Engineering(脑科学与认知工程系) Korea University(韩国大学) Dept. of Artificial Intelligence(人工智能系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15139 2025-11-20 q-bio.GN cs.AI cs.LG 74%

CASPER: Cross-modal Alignment of Spatial and single-cell Profiles for Expression Recovery

Amit Kumar, Maninder Kaur, Raghvendra Mall, Sukrit Gupta

机构 * Department of Computer Science \& Engineering, Indian Institute of Technology Ropar, India Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar. Department of Biomedical Engineering, Indian Institute of Technology Ropar, India

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14970 2025-11-20 cs.CV cs.AI cs.RO 62%

EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects

Gbenga Omotara, Ramy Farag, Seyed Mohamad Ali Tousi, G. N. DeSouza

机构 * Vision-Guided and Intelligent Robotics Lab (ViGIR) University of Missouri(视觉引导与智能机器人实验室(ViGIR)大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00275 2025-11-20 cs.CV cs.AI 62%

AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding

Md Asaduzzaman Jabin, Hanqi Jiang, Yiwei Li, Patrick Kaggwa, Eugene Douglass, Juliet N. Sekandi, Tianming Liu

机构 * University of Georgia(佐治亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 7th International Workshop on Large Scale Holistic Video Understanding: Toward Video Foundation Models

Journal ref Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07033 2025-11-20 cs.CV 57%

One Latent Space to Rule All Degradations: Unifying Restoration Knowledge for Image Fusion

Haolong Ma, Hui Li, Chunyang Cheng, Zeyang Zhang, Xiaoqing Luo, Xiaoning Song, Xiao-Jun Wu

机构 * Jiangnan University(江南大学) Suzhou University of Science and Technology(苏州科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15016 2025-11-20 cs.CV 57%

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

Zhenyu Cui, Jiahuan Zhou, Yuxin Peng

机构 * Zhenyu Cui, Jiahuan Zhou, Yuxin Peng(作者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05923 2025-11-20 cs.CV 57%

Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 57%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18757 2025-11-20 cs.CV 57%

ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * CCDS, NTU, Singapore(南洋理工大学新加坡分校) CCST, ZJUT, China(浙江工业大学计算机科学与技术学院) Terminus AI Lab, UCAS, China(中国科学院大学人工智能实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15251 2025-11-20 cs.LG cs.NI 50%

PLATONT: Learning a Platonic Representation for Unified Network Tomography

Chengze Du, Heng Xu, Zhiwei Yu, Bo Liu, Jialong Li

机构 * Computer Science and Control Engineering, Shenzhen University of Advanced Technology(深圳先进技术大学计算机科学与控制工程系) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14988 2025-11-20 cs.RO 50%

An Alignment-Based Approach to Learning Motions from Demonstrations

Alex Cuellar, Christopher K Fourie, Julie A Shah

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 8 pages, 8 figures, originally published in the IEEE Robotics and Automation Letters

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 11, pp. 11912-11919, Nov. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏