arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-15 至 2025-08-15 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2508.10371 2025-08-15 cs.RO 85%

Few-shot Vision-based Human Activity Recognition with MLLM-based Visual Reinforcement Learning

Wenqi Zheng, Yutaka Arakawa

机构 * Graduate School(研究生院) Faculty of Information Science(信息科学系) Electrical Engineering(电气工程) Kyushu University(九州大学) JAPAN(日本)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06293 2025-08-15 cs.CV 83%

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness

Qifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu, Bosheng Qin, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10552 2025-08-15 cs.CL cs.AI 81%

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Huyu Wu, Meng Tang, Xinhan Zheng, Haiyun Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10753 2025-08-15 cs.IR 78%

Hypercomplex Prompt-aware Multimodal Recommendation

Zheyu Chen, Jinfeng Xu, Hewei Wang, Shuo Yang, Zitong Wan, Haibo Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10696 2025-08-15 cs.CE 78%

Chem3DLLM: 3D Multimodal Large Language Models for Chemistry

Lei Jiang, Shuzhou Sun, Biqing Qi, Yuchen Fu, Xiaohua Xu, Yuqiang Li, Dongzhan Zhou, Tianfan Fu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10116 2025-08-15 cs.IR 67%

Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment

Yipeng Zhang, Hongju Yu, Aritra Mandal, Canran Xu, Qunzhi Zhou, Zhe Wu

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10704 2025-08-15 cs.CV 57%

Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios

Zhanwen Liu, Yujing Sun, Yang Wang, Nan Yang, Shengbo Eben Li, Xiangmo Zhao

机构 * School of Information Engineering, Chang’an University(信息工程学院,长安大学) School of Vehicle Mobility & College of AI, Tsinghua University(车辆运动学院与人工智能学院,清华大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10232 2025-08-15 cs.CV 57%

CellSymphony: Deciphering the molecular and phenotypic orchestration of cells with single-cell pathomics

Paul H. Acosta, Pingjun Chen, Simon P. Castillo, Maria Esther Salvatierra, Yinyin Yuan, Xiaoxi Pan

机构 * Translational Molecular Pathology Department, Division of Pathology and Laboratory Medicine, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心转化分子病理学部门) Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心肿瘤数据科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11049 2025-08-15 cs.LG cs.AI 57%

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning

Andrew P. Berg, Qian Zhang, Mia Y. Wang

机构 * Department of Computer Science(计算机科学系) College of Charleston(查尔斯顿学院) Department of Engineering(工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏