arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-31 至 2025-07-31 共收录 41 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2503.07631 2025-07-31 cs.LG cs.CL 57%

OWLViz: An Open-World Benchmark for Visual Question Answering

Thuy Nguyen, Dang Nguyen, Hoang Nguyen, Thuan Luong, Long Hoang Dang, Viet Dac Lai

机构 * Posts and Telecommunications Institute of Technology, Viet Nam(越南电信技术研究所) Adobe Research, USA(Adobe研究) University of Maryland, USA(美国马里兰大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments 8 pages + appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22389 2025-07-31 cs.RO cs.SY eess.SY 50%

Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators

Kaustav Chakraborty, Zeyuan Feng, Sushant Veer, Apoorva Sharma, Wenhao Ding, Sever Topan, Boris Ivanovic, Marco Pavone, Somil Bansal

机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) NVIDIA Research(NVIDIA研究)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 1 篇

2505.04390 2025-07-31 cond-mat.stat-mech 50%

Replica exchange nested sampling

Nico Unglert, Livia Bartók Pártay, Georg K. H. Madsen

专题命中 多模态Agent :multimodal(abstract)

Journal ref J. Chem. Theory Comput. (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 5 篇

2505.19010 2025-07-31 cs.CV cs.CL 84%

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

Md. Mithun Hossain, Md. Shakil Hossain, Sudipto Chaki, M. F. Mridha

机构 * Department of Computer Science and Engineering, Bangladesh University of Business and Technology(计算机科学与工程系,孟加拉国商业技术大学) Department of Computer Science, American International University-Bangladesh(计算机科学系,美国国际大学-孟加拉国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22880 2025-07-31 cs.IR 78%

AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS

Hai Ling, Tianchi Wang, Xiaohao Liu, Zhulin Tao, Lifang Yang, Xianglin Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

Comments 14 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22426 2025-07-31 cs.LG 78%

Multimodal Late Fusion Model for Problem-Solving Strategy Classification in a Machine Learning Game

Clemens Witt, Thiemo Leonhardt, Nadine Bergner, Mareen Grillenberger

机构 * TUD Dresden University of Technology(德累斯顿技术大学) RWTH Aachen University(亚琛工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments This is the author's version of a paper accepted for publication at the 2025 European Conference on Technology Enhanced Learning (EC-TEL 2025). The final authenticated version will be published in the Lecture Notes in Computer Science (LNCS) series by Springer and will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19442 2025-07-31 cs.AI cs.DC 70%

A Survey on Large Language Model Acceleration based on KV Cache Management

Haoyang Li, Yiming Li, Anxin Tian, Tianhao Tang, Zhanchao Xu, Xuejia Chen, Nicole Hu, Wei Dong, Qing Li, Lei Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Department of Computer Science and Engineering(计算机科学与工程系) Department of Computer Science and Technology(计算机科学与技术系) Department of Computing and Data Science(计算与数据科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025. The revised version incorporates more papers and has been further polished

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19582 2025-07-31 cs.RO 50%

An Actionable Hierarchical Scene Representation Enhancing Autonomous Inspection Missions in Unknown Environments

Vignesh Kottayam Viswanathan, Mario Alberto Valdes Saucedo, Sumeet Gajanan Satpute, Christoforos Kanellakis, George Nikolakopoulos

机构 * Robotics and AI(机器人与人工智能)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 3 篇

2507.22605 2025-07-31 physics.optics cond-mat.mes-hall physics.app-ph 78%

Optically Actuated Transitions in Multimodal, Bistable Micromechanical Oscillators

Lior Michaeli, Ramon Gao, Michael D. Kelzenberg, Claudio U. Hail, John E. Sader, Harry A. Atwater

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22464 2025-07-31 cs.LG cs.AI cs.MA stat.AP 57%

Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework

Peng-Yi Wu, Pei-Cing Huang, Ting-Yu Chen, Chantung Ku, Ming-Yen Lin, Yihuang Kang

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20300 2025-07-31 cs.HC cs.MM 57%

Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft

Xin Sun, Lei Wang, Yue Li, Jie Li, Massimo Poesio, Julian Frommel, Koen Hinriks, Jiahuan Pei

专题命中 其他多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏