arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-06 至 2025-08-06 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 5 篇

2508.03102 2025-08-06 cs.CV 83%

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

Tianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01455 2025-08-06 cs.CV cs.AI cs.LG 81%

Automatic Fused Multimodal Deep Learning for Plant Identification

Alfreds Lapkovskis, Natalia Nefedova, Ali Beikmohammadi

机构 * Department of Computer and Systems Sciences(计算机与系统科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Front. Plant Sci., 05 August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08344 2025-08-06 cs.CV 79%

MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

Jihao Gu, Fei Wang, Kun Li, Yanyan Wei, Zhiliang Wu, Dan Guo

机构 * University College London (UCL)(伦敦大学学院) School of Computer Science(计算机科学学院) Information Engineering, School of Artificial Intelligence, Hefei University of Technology (HFUT)(信息工程学院,人工智能学院,合肥工业大学) ReLER, CCAI, Zhejiang University, China(ReLER、CCAI、浙江大学,中国) Key Laboratory of Knowledge Engineering with Big Data (HFUT), Ministry of Education(大数据知识工程重点实验室(HFUT),教育部) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国) Xinsight Lab, Research Institute, Hefei Zhongjuyuan Intelligent Technology Co., Ltd., China(Xinsight实验室,研究院,合肥中睿智能科技有限公司,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 1st Place in Micro-gesture Classification sub-challenge in 3rd MiGA at IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02562 2025-08-06 cs.CV cs.AI cs.CY cs.LG 62%

Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification

Prateek Mittal, Puneet Goyal, Joohi Chauhan

机构 * Motilal Nehru National Institute of Technology Allahabad, India(Motilal Nehru国立技术学院Allahabad分校) Indian Institute of Technology Ropar, India(印度理工学院Ropar分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03277 2025-08-06 cs.CV 57%

Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification

Hang Guo, Qing Zhang, Zixuan Gao, Siyuan Yang, Shulin Peng, Xiang Tao, Ting Yu, Yan Wang, Qingli Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACMMM'25

详情

展开后加载摘要…

URL PDF HTML 收藏