arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-09 至 2025-10-09 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2505.02152 2025-10-09 cs.RO 82%

Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions

Cunxin Fan, Xiaosong Jia, Yihang Sun, Yixiao Wang, Jianglan Wei, Ziyang Gong, Xiangyu Zhao, Masayoshi Tomizuka, Xue Yang, Junchi Yan, Mingyu Ding

机构 * Shanghai Jiao Tong University(上海交通大学) UC Berkeley(伯克利大学) UNC, Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03363 2025-10-09 cs.CV cs.AI eess.IV 62%

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang, Lingqiao Liu, Dacheng Tao, Tianyou Chai, Xiatian Zhu

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) University of Surrey(Surrey大学) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18269 2025-10-09 cs.CV cs.AI 62%

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

Sangho Lee, Il Yong Chun, Hogun Park

机构 * Sangho Lee 1,2(Sangho Lee 教授) Il Yong Chun 1,3(Il Yong Chun 教授) Hogun Park 1(Hogun Park 教授)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to the AAAI 2025 Main Technical Track. This is an extended version of the original submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07098 2025-10-09 cs.CL 57%

TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription

Guo Yutong, Wanying Wang, Yue Wu, Zichen Miao, Haoyu Wang

机构 * Johns Hopkins University(约翰霍普金斯大学) Purdue University(普渡大学) University at Albany(阿尔巴尼大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏