arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-02 至 2025-12-02 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2512.00496 2025-12-02 cs.CL cs.AI 87%

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

CACARA:基于文本中心方法的跨模态对齐,用于高效多模态和多语言学习

Diego A. B. Moreira, Alef I. Ferreira, Jhessica Silva, Gabriel O. dos Santos, Gustavo Bonil, João Gondim, Marina dos Santos, Helena Maia, Simone Hashiguti, Nádia da Silva, Carolina Scarton, Helio Pedrini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP), Brasil(计算机学院,Campinas州立大学(UNICAMP)) Instituto de Estudos da Linguagem, Universidade Estadual de Campinas (UNICAMP), Brasil(语言研究学院,Campinas州立大学(UNICAMP)) Instituto de Informática, Universidade Federal de Goiás (UFG), Goiás, Brasil(信息学院,戈亚斯联邦大学(UFG)) Department of Computer Science, University of Sheffield, Sheffield, United Kingdom(计算机科学系,谢菲尔德大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title);分类 cs.CL、cs.AI

AI总结 CACARA通过文本中心方法实现多模态和多语言学习,无需重新训练即可支持100多种语言,提升音频到文本检索性能达14.24个百分点。

Comments 25 pages, 12 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00363 2025-12-02 cs.CV 83%

MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters

MM-DETR: 一种高效的多模态检测Transformer,采用Mamba驱动的双粒度融合和频率感知模态适配器

Jianhong Han, Yupei Wang, Yuan Zhang, Liang Chen

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Beijing Institute of Technology Chongqing Innovation Center(北京理工大学重庆创新中心) National Key Laboratory for Space-Born Intelligent Information Processing(空间智能信息处理国家级重点实验室) School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MM-DETR通过Mamba驱动的双粒度融合和频率感知模态适配器,实现高效的多模态目标检测,提升检测精度与轻量化性能。

Comments Manuscript submitted to IEEE Transactions on Geoscience and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01442 2025-12-02 cs.MM cs.AI 81%

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

PSA-MF:基于人格-情感对齐的多级融合用于多模态情感分析

Heng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao, Xuefei Liu, Ruibo Fu, Changsheng Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 PSA-MF通过引入人格-情感对齐和多级融合方法,提升多模态情感分析的识别性能。

Comments AAAI 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01214 2025-12-02 cs.CV cs.AI 81%

M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis

M4-BLIP:通过面部增强的局部分析推进多模态媒体篡改检测

Hang Wu, Ke Sun, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing(多媒体可信感知与高效计算重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 M4-BLIP通过引入面部增强的局部分析,提升多模态媒体篡改检测的准确性和可解释性。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08679 2025-12-02 cs.CV cs.AI 81%

MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion

MMIF-AMIN: 适应性损失驱动的多尺度可逆密集网络用于多模态医学图像融合

Tao Luo, Weihua Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 MMIF-AMIN通过可逆密集网络和多尺度互补特征提取模块,实现多模态医学图像融合的高效融合与精准诊断。

Comments This manuscript is withdrawn to allow for substantial expansion and restructuring. Based on recent research progress, we plan to add Generalization experiment and reorganize the manuscript structure to improve readability and logical flow. Thank you for your understanding and support

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01949 2025-12-02 cs.CV 79%

Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models

脚本:图结构和查询条件的语义令牌修剪用于多模态大语言模型

Zhongyu Yang, Dannong Xu, Wei Pang, Yingfang Yuan

机构 * BCML, Heriot-Watt University(赫瑞瓦德大学BCML中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 Script通过图结构和查询条件的语义令牌修剪,提升多模态大语言模型的效率和准确性,实现显著的性能提升。

Comments Published in Transactions on Machine Learning Research, Project in https://01yzzyu.github.io/script.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20469 2025-12-02 q-bio.QM cs.CV 79%

Prediction of Distant Metastasis in Head and Neck Cancer Patients Using Tumor and Peritumoral Multi-Modal Deep Learning

利用肿瘤及周围多模态深度学习预测头颈癌患者远端转移

Nuo Tong, Changhao Liu, Zizhao Tang, Feifan Sun, Yingping Li, Shuiping Gou, Mei Shi

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本研究提出多模态深度学习模型,结合CT影像、放射组学和临床数据,用于预测头颈癌患者远端转移风险,通过多模态融合显著提高预测性能。

Comments 23 pages, 6 figures, 7 tables. Nuo Tong and Changhao Liu contributed equally. Corresponding Authors: Shuiping Gou and Mei Shi

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01410 2025-12-02 cs.CL 79%

DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis

DyFuLM:一种用于情感分析的先进多模态框架

Ruohan Zhou, Jiachen Yuan, Churui Yang, Wenzheng Huang, Guoyan Zhang, Shiyao Wei, Jiazhen Hu, Ning Xin, Md Maruf Hasan

机构 * Department of Applied Mathematics, Xi'an Jiaotong-Liverpool University(应用数学系,西安交通大学-利物浦大学) School of AI and Advanced Computing, XJTLU Entrepreneur College (Taicang)(人工智能与先进计算学院,XJTLU创业学院(太仓)) Department of Intelligent Science, Xi'an Jiaotong-Liverpool University(智能科学系,西安交通大学-利物浦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 DyFuLM通过动态融合和门控聚合模块提升多模态情感分析的准确率与稳定性

Comments 8 pages, 6 figures, preprint. Under review for a suitable AI conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00596 2025-12-02 cs.IR cs.AI 79%

DLRREC: Denoising Latent Representations via Multi-Modal Knowledge Fusion in Deep Recommender Systems

DLRREC: 通过深度融合多模态知识在深度推荐系统中进行潜在表示去噪

Jiahao Tian, Zhenkai Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 DLRREC通过深度融合多模态和协同知识,提升深度推荐系统中潜在表示的去噪能力,从而实现更精确的推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01750 2025-12-02 eess.SP cs.LG 78%

Multimodal Mixture-of-Experts for ISAC in Low-Altitude Wireless Networks

多模态专家混合模型用于低空无线网络中的ISAC

Kai Zhang, Wentao Yu, Hengtao He, Shenghui Song, Jun Zhang, Khaled B. Letaief

机构 * IEEE

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种多模态专家混合模型,用于提升低空无线网络中ISAC的性能,通过自适应融合策略提高环境感知和通信效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00042 2025-12-02 cs.CV cs.AI cs.CL cs.CY 67%

Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions

弥合差距:面向标准化考试题目的视觉语言模型数据驱动微调

Egemen Sert, Şeyda Ertekin

机构 * organization= Department of Computer Engineering, Middle East Technical University (METU) , city= Ankara , country= Türkiye organization= METU-DTX Digital Transformation \& Innovation Centre, METU , city= Ankara , country= Türkiye

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究通过高质量数据和优化语法提升视觉语言模型在标准化考试题目的多模态推理性能,达到接近SOTA的水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10194 2025-12-02 cs.CV 57%

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07062 2025-12-02 cs.AI 57%

Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision

通过噪声长描述监督提升城市影像区域表示学习

Yimei Zhang, Guojiang Shen, Kaili Ning, Tongwei Ren, Xuebo Qiu, Mengmeng Wang, Xiangjie Kong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出UrbanLN框架,通过长文本意识和噪声抑制提升城市影像区域表示学习,有效解决细粒度特征对齐与噪声干扰问题。

Comments Accepted as a full paper by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00293 2025-12-02 cs.LG cs.AI 57%

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

FiCoTS: 细到粗的LLM增强层次跨模态交互用于时间序列预测

Yafei Lyu, Hao Zhou, Lu Zhang, Xu Yang, Zhiyong Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所MAIS) Great Bay University(大亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 FiCoTS通过细到粗的LLM增强层次跨模态交互框架,提升多模态时间序列预测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01358 2025-12-02 cs.RO cs.LG 50%

Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1

模态增强的基座机器人策略微调用于GR1和G1跨躯体操作

Junsung Park, Hogun Kee, Songhwai Oh

机构 * Department of Electrical and Computer Engineering, Seoul National University(电气与计算机工程系,首尔国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出了一种模态增强的微调方法,通过多模态数据提升机器人策略在不同躯体上的性能,显著提高了任务成功率。

Comments 8 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00063 2025-12-02 q-bio.NC 50%

Exploring the changes in brain network SC-FC coupling patterns of partial sleep deprivation based on DTI-fMRI fusion analysis

基于DTI-fMRI融合分析探讨部分睡眠剥夺对脑网络SC-FC耦合模式变化的影响

Mengyuan Liu, Jing Hu, Zhenzhen Ru, Ruomeng Quan, Xu Zhang, Ning Qiang, Jin Li

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本研究通过DTI-fMRI融合分析,探讨部分睡眠剥夺对脑网络SC-FC耦合模式的影响,发现睡眠剥夺导致神经网络结构和功能连接的异常,提出SC-FC耦合方法作为睡眠相关情绪失调的新型生物标记物。

详情

展开后加载摘要…

URL PDF HTML 收藏