arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-31 至 2025-07-31 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 5 篇

2505.19010 2025-07-31 cs.CV cs.CL 84%

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

Md. Mithun Hossain, Md. Shakil Hossain, Sudipto Chaki, M. F. Mridha

机构 * Department of Computer Science and Engineering, Bangladesh University of Business and Technology(计算机科学与工程系,孟加拉国商业技术大学) Department of Computer Science, American International University-Bangladesh(计算机科学系,美国国际大学-孟加拉国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22880 2025-07-31 cs.IR 78%

AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS

Hai Ling, Tianchi Wang, Xiaohao Liu, Zhulin Tao, Lifang Yang, Xianglin Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

Comments 14 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22426 2025-07-31 cs.LG 78%

Multimodal Late Fusion Model for Problem-Solving Strategy Classification in a Machine Learning Game

Clemens Witt, Thiemo Leonhardt, Nadine Bergner, Mareen Grillenberger

机构 * TUD Dresden University of Technology(德累斯顿技术大学) RWTH Aachen University(亚琛工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments This is the author's version of a paper accepted for publication at the 2025 European Conference on Technology Enhanced Learning (EC-TEL 2025). The final authenticated version will be published in the Lecture Notes in Computer Science (LNCS) series by Springer and will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19442 2025-07-31 cs.AI cs.DC 70%

A Survey on Large Language Model Acceleration based on KV Cache Management

Haoyang Li, Yiming Li, Anxin Tian, Tianhao Tang, Zhanchao Xu, Xuejia Chen, Nicole Hu, Wei Dong, Qing Li, Lei Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Department of Computer Science and Engineering(计算机科学与工程系) Department of Computer Science and Technology(计算机科学与技术系) Department of Computing and Data Science(计算与数据科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025. The revised version incorporates more papers and has been further polished

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19582 2025-07-31 cs.RO 50%

An Actionable Hierarchical Scene Representation Enhancing Autonomous Inspection Missions in Unknown Environments

Vignesh Kottayam Viswanathan, Mario Alberto Valdes Saucedo, Sumeet Gajanan Satpute, Christoforos Kanellakis, George Nikolakopoulos

机构 * Robotics and AI(机器人与人工智能)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏