arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2508.05612 2026-03-04 cs.LG cs.AI 83%

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Shuffle-R1: 通过数据导向的动态洗牌提升多模态大语言模型的强化学习框架

Linghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Bin Qin, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(MiLM Plus,小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 Shuffle-R1通过动态洗牌和轨迹采样提升多模态大语言模型的强化学习效率,实现更高效的训练效果。

Comments This paper has been accepted by ICLR 2026. Conference link: https://iclr.cc/virtual/2026/poster/10007559 OpenReview link: https://openreview.net/forum?id=mYP33u1QBK Project page at: https://xenozlh.github.io/Shuffle-R1/

Journal ref The Fourteenth International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01784 2026-03-03 cs.CR cs.AI 83%

Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution

基于结构对抗进化的多模态对齐

Guoxin Shi, Haoyu Wang, Zaihui Yang, Yuxing Wang, Yongzhe Chang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 本文提出CEMA框架,通过共进化对抗提升多模态对齐的鲁棒性和安全性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00482 2026-03-03 cs.CV cs.IT math.IT 83%

TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications

TokenCom: 一种用于多模态和多任务令牌通信的视觉-语言模型

Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li, Kezhi Wang, Cunhua Pan, Zhu Han, Jiangzhou Wang

机构 * Hunan Provincial Key Laboratory of Intelligent Computing and Language Information Processing, Hunan Normal University(湖南省级智能计算与语言信息处理重点实验室,湖南师范大学) School of Information Science and Engineering, Hunan Normal University(信息科学与工程学院,湖南师范大学) Changsha Social Laboratory of Artificial Intelligence, Hunan University of Technology and Business(长沙人工智能社会实验室,湖南工业大学) School of Computer Science, Hunan University of Technology and Business(计算机科学学院,湖南工业大学) Department of Computer Science, Brunel University London(伦敦布鲁内尔大学计算机科学系) National Mobile Communications Research Laboratory, Southeast University(东南大学国家移动通信研究中心) Department of Electrical and Computer Engineering, University of Houston(电子与计算机工程系,休斯顿大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 TokenCom提出了一种新的视觉-语言模型框架TaiChi,通过双视觉分词器和双向注意力网络提升多模态和多任务令牌通信的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05394 2026-03-02 cs.CV cs.LG 83%

pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models

pFedMMA: 为视觉-语言模型设计的个性化联邦微调多模态适配器

Sajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 pFedMMA通过多模态适配器实现视觉-语言模型的个性化联邦微调,平衡个性化与泛化能力,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22917 2026-02-27 cs.CV 83%

Towards Multimodal Domain Generalization with Few Labels

迈向少标签多模态领域泛化的研究

Hongzhao Li, Hao Dong, Hualei Wan, Shupan Li, Mingliang Xu, Muhammad Haris Khan

机构 * Zhengzhou University(郑州大学) ETH Zürich(苏黎世联邦理工学院) MBZUAI(马克斯·普朗克智能系统研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出SSMDG框架,通过三个关键组件实现少标签多模态领域泛化,提升跨模态鲁棒性和领域不变性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20818 2026-02-25 cs.CV 83%

GatedCLIP: Gated Multimodal Fusion for Hateful Memes Detection

GatedCLIP:用于仇恨表情包检测的门控多模态融合

Yingying Guo, Ke Zhang, Zirong Zeng

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 GatedCLIP通过门控融合和对比学习提升多模态仇恨表情包检测性能,实现0.66 AUROC优于CLIP基线。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20046 2026-02-24 cs.CV cs.LG 83%

Closing the gap in multimodal medical representation alignment

弥合多模态医学表征对齐的差距

Eleonora Grassucci, Giordano Cicchetti, Danilo Comminiello

机构 * Dept. of Information Engineering, Electronics, and Telecommunications(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种模态无关的框架,用于弥合多模态医学表征对齐中的模态差距,提升放射学图像与临床文本之间的对齐效果。

Comments Accepted at MLSP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19832 2026-02-24 cs.CV 83%

M3S-Net: Multimodal Feature Fusion Network Based on Multi-scale Data for Ultra-short-term PV Power Forecasting

M3S-Net:基于多尺度数据的多模态特征融合网络用于超短期光伏功率预测

Penghui Niu, Taotao Cai, Suqi Zhang, Junhua Gu, Ping Zhang, Qiqi Liu, Jianxin Li

机构 * School of Artificial Intelligence, Hebei University of Technology(人工智能学院,河北工业大学) University of Southern Queensland(南方昆士兰大学) School of Information Engineering, Tianjin University of Commerce(信息工程学院,天津商业大学) Hebei Province Key Laboratory of Big Data Calculation, Hebei University of Technology(大数据计算河北重点实验室,河北工业大学) General AI Lab, School of Engineering, Westlake University(通用人工智能实验室,西湖大学) School of Business and Law, Edith Cowan University(商学院和法学院,埃迪斯科文大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 M3S-Net通过多尺度数据融合和跨模态Mamba交互模块,提升超短期光伏功率预测的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19505 2026-02-24 cs.CV 83%

Test-Time Computing for Referring Multimodal Large Language Models

测试时计算用于指代多模态大语言模型

Mingrui Wu, Hao Chen, Jiayi Ji, Xiaoshuai Sun, Zhiyuan Liu, Liujuan Cao, Ming-Ming Cheng, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) VCIP, CS, Nankai University(VCIP,计算机科学,南开大学) Tsinghua University(清华大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 ControlMLLM++通过注入可学习的视觉提示实现多模态大语言模型的测试时适应,提升细粒度视觉推理能力。

Comments arXiv admin note: substantial text overlap with arXiv:2407.21534

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17196 2026-02-20 cs.CV 83%

EntropyPrune: Matrix Entropy Guided Visual Token Pruning for Multimodal Large Language Models

EntropyPrune: 基于矩阵熵的视觉令牌修剪用于多模态大语言模型

Yahong Wang, Juncheng Wu, Zhangkai Ni, Chengmei Yang, Yihang Liu, Longzhen Yang, Yuyin Zhou, Ying Wen, Lianghua He

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) University of California, Santa Cruz(加州大学圣克ruz分校) East China Normal University(华东师范大学) Shanghai Eye Disease Prevention and Treatment Center(上海眼病预防与治疗中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 EntropyPrune通过矩阵熵指导的视觉令牌修剪方法,提升多模态大语言模型的推理效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08126 2026-02-13 cs.CV 83%

MambaFusion: Adaptive State-Space Fusion for Multimodal 3D Object Detection

MambaFusion:多模态3D目标检测的自适应状态空间融合

Venkatraman Narayanan, Bala Sai, Rahul Ahuja, Pratik Likhar, Varun Ravi Kumar, Senthil Yogamani

机构 * Qualcomm Technologies, Inc(高通技术公司) Qualcomm India Private Limited(高通印度私人有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 MambaFusion通过结合选择性状态空间模型与窗口化变换器,实现高效且可靠的多模态3D目标检测,提升自动驾驶系统的感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22463 2026-02-12 cs.MM 83%

Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation

正交解缠与投影特征对齐用于对话中多模态情绪识别

Xinyi Che, Wenbo Wang, Jian Guan, Qijun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出OD-PFA框架,通过正交解缠与投影特征对齐技术提升对话中多模态情绪识别性能。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05885 2026-02-09 cs.IR cs.AI 83%

An item is worth one token in Multimodal Large Language Models-based Sequential Recommendation

在基于多模态大语言模型的序列推荐中,一个项目相当于一个标记

Qiyong Zhong, Jiajie Su, Ming Yang, Yunshan Ma, Xiaolin Zheng, Chaochao Chen

机构 * Zhejiang University(浙江大学) Singapore Management University(新加坡国立管理学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 Speeder通过多模态表示压缩、模态感知渐进优化和序列位置感知增强,提升多模态大语言模型在序列推荐中的效率与效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05729 2026-02-06 cs.CV cs.LG 83%

Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification

自适应全局与细粒度感知融合用于兼容硬负样本放大的人脸嵌入

Lexiang Hu, Youze Xue, Dian Li, Gang Liu, Zhouchen Lin

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(人工智能国家重点实验室,智能科学与技术学院,北京大学) Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出AGFF-Embed方法,通过自适应融合全局和细粒度语义信息,提升多模态嵌入在一般和细粒度理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22936 2026-02-06 cs.CV 83%

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

PPE:用于多模态大语言模型中token压缩的位置保持嵌入

Mouxiao Huang, Borui Jiang, Dehua Zheng, Hailin Hu, Kai Han, Xinghao Chen

机构 * Huawei Technologies(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 PPE通过保持位置信息提升多模态大语言模型的token压缩效率和性能

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15098 2026-02-04 cs.LG cs.AI 83%

How Intermodal Interaction Affects the Performance of Deep Multimodal Fusion for Mixed-Type Time Series

多模态交互如何影响深度多模态融合在混合类型时间序列中的性能

Simon Dietz, Thomas Altstidl, Dario Zanca, Björn Eskofier, An Nguyen

机构 * FAU Erlangen-Nürnberg(弗赖堡大学埃尔朗根-纽伦堡分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文研究了多模态交互对深度多模态融合在混合类型时间序列预测中的影响,通过三种融合类型和五种融合方法的比较,揭示了交互强度和方向对融合策略选择的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02784 2026-02-04 cs.LG cs.AI 83%

Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning

跨时间注意力融合(CTAF)用于自监督学习中的多模态生理信号

Arian Khorasani, Théophile Demazure

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 CTAF通过时间感知的交叉注意力机制,实现多模态生理信号在自监督学习中的高效融合,提升匹配对的余弦边距和跨模态检索性能,同时保持高准确率并减少标签依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21280 2026-02-02 cs.CV 83%

Token Entropy Regularization for Multi-modal Antenna Affiliation Identification

基于令牌熵正则化的多模态天线归属识别

Dong Chen, Ruoyu Li, Xinyan Zhang, Jialei Xu, Ruosen Zhao, Zhikang Zhang, Lingyun Li, Zizhuang Wei

机构 * Huawei(华为) The University of Hong Kong(香港大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于令牌熵正则化的多模态天线归属识别方法,通过融合视频、几何特征和PCI信号,提升通信网络优化效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22020 2026-01-30 cs.LG cs.CV 83%

Visual-Guided Key-Token Regularization for Multimodal Large Language Model Unlearning

多模态大语言模型去敏中的视觉引导关键标记正则化

Chengyi Cai, Zesheng Ye, Peike Li, Bo Han, Jianzhong Qi, Feng Liu

机构 * The University of Melbourne(墨尔本大学) Google Research(谷歌研究) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出视觉引导的关键标记正则化方法,用于多模态大语言模型的去敏,通过信息熵定义关键标记并利用梯度重新加权提升去敏效果,实验表明能有效减少遗忘并保持响应一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22009 2026-01-30 cond-mat.mtrl-sci cs.AI cs.LG physics.comp-ph 83%

MEIDNet: Multimodal generative AI framework for inverse materials design

MEIDNet: 多模态生成式AI框架用于反向材料设计

Anand Babu, Rogério Almeida Gouvêa, Pierre Vandergheynst, Gian-Marco Rignanese

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MEIDNet通过多模态生成式AI框架实现反向材料设计,利用对比学习提升学习效率,生成低带隙钙钛矿结构并验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21547 2026-01-30 cs.LG cs.AI 83%

Multi-Modal Time Series Prediction via Mixture of Modulated Experts

多模态时间序列预测 via 专家混合

Lige Zhang, Ali Maatouk, Jialin Chen, Leandros Tassiulas, Rex Ying

机构 * Duke Kunshan University, China(杜克昆山大学) Yale University, USA(耶鲁大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出专家调制方法,通过条件控制专家行为实现多模态时间序列预测的高效跨模态控制。

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20258 2026-01-29 cs.CV cs.LG 83%

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

模态平衡的协同蒸馏用于多模态领域泛化

Xiaohan Wang, Zhangtao Cheng, Ting Zhong, Leiting Chen, Fan Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MBCD框架,通过自适应模态丢弃、梯度一致性约束和跨模态蒸馏,解决多模态领域泛化中模态不平衡问题,提升模型泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15139 2026-01-27 cs.CV 83%

Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis

统一的跨模态注意力-混合器基于结构-功能连接组融合的神经精神疾病诊断

Badhan Mazumder, Lei Wu, Vince D. Calhoun, Dong Hye Ye

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University(跨机构神经影像与数据科学转化研究中心(TReNDS),佐治亚州立大学、佐治亚理工学院和埃默里大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出ConneX方法,通过统一的跨模态注意力和MLP-Mixer实现结构-功能连接组的多模态融合,提升神经精神疾病诊断性能。

Comments Published in the Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025). IEEE Xplore. DOI: 10.1109/EMBC58623.2025.11254194

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16381 2026-01-26 cs.CV 83%

VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection

VTFusion: 一种面向少样本异常检测的视觉-文本多模态融合网络

Yuxin Jiang, Yunkang Cao, Yuqi Cheng, Yiheng Zhang, Weiming Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 VTFusion通过自适应特征提取和多模态融合模块,提升少样本异常检测性能,实现96.8%的AUROC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14274 2026-01-22 cs.LG cs.AI 83%

Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation

分割与精炼:提升对话中情感识别的多模态表示与可解释性

Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le

机构 * VNU University of Engineering and Technology(越南工程大学) FPT Software AI Center(FPT软件人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出DnR框架,通过分割和精炼多模态表示提升对话中情感识别的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02438 2026-01-21 cs.SE cs.AI cs.CR 83%

Focus on What Matters: Fisher-Guided Adaptive Multimodal Fusion for Vulnerability Detection

聚焦关键要素:基于Fisher信息的自适应多模态融合用于漏洞检测

Yun Bian, Yi Chen, HaiQuan Wang, ShiHao Li, Zhe Cui

机构 * University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出TaCCS-DFA框架,通过Fisher信息引导的自适应多模态融合,提升漏洞检测的F1分数,同时降低推理延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05097 2026-01-14 cs.CL 83%

MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification

MultiCheck: 通过统一多模态事实验证加强网络信任

Aditya Kishore, Gaurav Kumar, Jasabanta Patro

机构 * IISER Bhopal(比哈尔IISER)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 MultiCheck通过统一多模态事实验证框架,提升网络信任,具备高效、透明和抗噪能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05538 2026-01-12 cs.CV 83%

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

DIFF-MF: 一种基于差分驱动的通道-空间状态空间模型用于多模态图像融合

Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu

机构 * School of Automation, Southeast University(东南大学自动化学院) Low-Altitude Intelligence Lab, Xiong’an National Innovation Center Technology Co., Ltd.(雄安国家创新中心技术有限公司低空智能实验室) Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创蓝天科技有限公司) School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 DIFF-MF通过差分驱动的通道-空间状态空间模型,有效整合多模态图像信息,提升融合图像的质量和显著性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02249 2026-01-06 cs.CV 83%

SLGNet: Synergizing Structural Priors and Language-Guided Modulation for Multimodal Object Detection

SLGNet: 结合结构先验与语言引导调制的多模态目标检测

Xiantai Xiang, Guangyao Zhou, Zixiao Wen, Wenshuai Li, Ben Niu, Feng Wang, Lijia Huang, Qiantong Wang, Yuhan Liu, Zongxu Pan, Yuxin Hu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所) Key Laboratory of Target Cognition and Application Technology, Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 SLGNet通过结合结构先验与语言引导调制,在冻结的ViT基础上实现高效多模态目标检测,提升环境适应性和检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01339 2026-01-06 cs.CV 83%

Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning

通过脑启发的分层表征学习实现细粒度跨模态理解

Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Tianming Liu, Fei Dou

机构 * School of Computing, University of Georgia, Athens, GA, USA(计算学院,佐治亚大学,亚特兰大,GA,USA)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 NeuroAlign通过脑启发的分层表征学习,在fMRI-视频对齐中实现细粒度跨模态理解,提升跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏