arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-16 至 2025-10-16 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2504.02477 2025-10-16 cs.RO cs.CV 83%

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng, Lue Fan, Dong An, Changwei Wang, Li Guo, Weiliang Meng, Xiaopeng Zhang, Rongtao Xu, Shibiao Xu

机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts Telecommunications, China [1ex] dKey Laboratory of Computing Power Network Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet Service Computing, Shandong Fundamental Research Center for Computer Science, China

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026

Journal ref Information Fusion, 126 (Part B), February 2026, 103652

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13131 2025-10-16 cs.CV cs.MM 81%

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

Rongjun Chen, Chengsi Yao, Jinchang Ren, Xianxian Zeng, Peixian Wang, Jun Yuan, Jiawen Li, Huimin Zhao, Xu Lu

机构 * School of Computer Science, Guangdong Polytechnic Normal University(广东 polytechnic 正规大学计算机学院)

专题命中 多模态训练与对齐 :image-text(title);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13076 2025-10-16 cs.CV 74%

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

Hao Zhou, Zhanning Gao, Zhili Chen, Maosheng Ye, Qifeng Chen, Tongyi Cao, Honggang Qi

机构 * University of Chinese Academy of Sciences(中国科学院大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04247 2025-10-16 cs.IR cs.MM 70%

I$^3$-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality Recommendation

Huilin Chen, Miaomiao Cai, Fan Liu, Zhiyong Cheng, Richang Hong, Meng Wang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.MM

Comments ACM Multimedia 2025 Accepted

Journal ref In Proceedings of the 33st ACM International Conference on Multimedia (MM '25), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21976 2025-10-16 cs.CV cs.AI 62%

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, Xiang Li

机构 * College of Computer Science and Technology of Zhejiang University(浙江大学计算机科学与技术学院) Polytechnic Institute of Zhejiang University(浙江大学Polytechnic学院) Om AI Research(Om AI研究机构) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究机构) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) School of Mathematical Sciences of Zhejiang University(浙江大学数学科学学院) China Academy of Space Technology(中国航天科技研究院) University of Bristol(布里斯托大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13316 2025-10-16 cs.CV 57%

Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests

Fitim Abdullahu, Helmut Grabner

机构 * Zurich University of Applied Sciences(苏黎世应用科学大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08163 2025-10-16 cs.CL 57%

ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code

Jian Xie, Zhendong Chu, Aoxiao Zhong, Kai Zhang, Mingzhe Han, Xing Fan, Jialie Shen, Qingsong Wen

机构 * Squirrel Ai Learning The Ohio State University(俄亥俄州立大学) City St George’s, University of London(伦敦城市圣乔治学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07613 2025-10-16 cs.CV 57%

Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease

Fangqi Cheng, Surajit Ray, Xiaochen Yang

机构 * School of Mathematics and Statistics, University of Glasgow, UK(数学与统计学学院,格拉斯哥大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted at MICAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12971 2025-10-16 cs.RO 50%

Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation

Anran Zhang, Hanzhi Chen, Yannick Burkhardt, Yao Zhong, Johannes Betz, Helen Oleynikova, Stefan Leutenegger

机构 * ETH Zurich(苏黎世联邦理工学院) Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏