arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-15 至 2025-10-15 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 7 篇

2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12299 2025-10-15 cs.IR 67%

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu, Julio Vizcarra, Masato Taya

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16845 2025-10-15 cs.CV cs.AI cs.LG 62%

NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows

Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Nikita Lyubaykin, Andrei Polubarov, Alexander Derevyagin, Vladislav Kurenkov

机构 * ETH Zürich(苏黎世联邦理工学院) MIPT(莫斯科国立信息安全大学) Skoltech(斯克里普斯基技术大学) Innopolis University(因诺波利斯大学) HSE(俄罗斯高等经济学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments https://github.com/dunnolab/NinA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12483 2025-10-15 cs.RO cs.CV 57%

Fast Visuomotor Policy for Robotic Manipulation

Jingkai Jia, Tong Yang, Xueyao Chen, Chenhuan Liu, Wenqiang Zhang

机构 * Fudan University(复旦大学) MEGVII Technology(MEGVII科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12185 2025-10-15 cs.CL cs.SD 57%

Not in Sync: Unveiling Temporal Bias in Audio Chat Models

Jiayu Yao, Shenghua Liu, Yiwei Wang, Rundong Cheng, Lingrui Mei, Baolong Bi, Zhen Xiong, Xueqi Cheng

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学默塞德分校) Beijing University of Posts and Telecommunications(北京邮电大学) University of Southern California(南加州大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07984 2025-10-15 cs.CV 57%

OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding

Jingli Lin, Chenming Zhu, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 30 pages, a benchmark designed to evaluate Online Spatio-Temporal understanding from the perspective of an agent actively exploring a scene. Project Page: https://rbler1234.github.io/OSTBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12434 2025-10-15 cs.CV 57%

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

Qi Wang, Yanrui Yu, Ye Yuan, Rui Mao, Tianfei Zhou

机构 * Beijing Institute of Technology(北京理工大学) Shenzhen University(深圳大学)

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. Code: https://github.com/QiWang98/VideoRFT

详情

展开后加载摘要…

URL PDF HTML 收藏