arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-08 至 2025-12-08 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 4 篇

2512.05515 2025-12-08 cs.CV cs.LG 83%

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

DashFusion: 基于分层瓶颈融合的双流对齐多模态情感分析

Yuhua Wen, Qifei Li, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学信息科学与技术国家研究中心) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 DashFusion通过双流对齐与分层瓶颈融合技术,提升多模态情感分析的性能与效率。

Comments Accepted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16054 2025-12-08 cs.CV 83%

Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model

基于多模态大语言模型的语言引导推理用于群体活动检测

Jihua Peng, Qianxiong Xu, Yichen Liu, Chenxi Liu, Cheng Long, Rui Zhao, Ziyue Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出LIR-GAD框架,通过多模态大语言模型实现群体活动检测,引入活动标记和群体标记以提升语义理解和分类性能。

Comments This work is being incorporated into a larger study

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05481 2025-12-08 cs.CV cs.AI 62%

UniFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial Fusion

UniFS: 通过频率-空间融合实现统一的多对比MRI重建

Jialin Li, Yiwei Ren, Kai Pan, Dong Wei, Pujin Cheng, Xian Wu, Xiaoying Tang

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China(电子与电气工程系,南方科技大学,深圳,中国) Jarvis Research Center, Tencent YouTu Lab, China(Jarvis研究中心,腾讯YouTu实验室,中国) Department of Electrical and Electronic Engineering, University of Hong Kong, Hong Kong, China(电子与电气工程系,香港大学,香港,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 UniFS通过频率-空间融合模块实现多对比MRI重建的统一处理,提升模型泛化能力,适用于多种k空间欠采样模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05593 2025-12-08 cs.CV 57%

Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer

通过无骨骼绑定图像传输学习高保真布料动画

Rong Wang, Wei Mao, Changsheng Lu, Hongdong Li

机构 * The Australian National University(澳大利亚国立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本研究提出无骨骼绑定图像传输方法,通过独立估计顶点位置和法线以生成高保真布料动画,提升动画质量和细节恢复能力。

Comments Accepted to 3DV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏