arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2511.06281 2025-11-11 cs.CV 57%

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang, Daizong Liu, Yu Cheng

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05636 2025-11-10 cs.LG cs.AI 57%

Graph Learning

Feng Xia, Ciyuan Peng, Jing Ren, Falih Gozi Febrinanto, Renqiang Luo, Vidya Saikrishna, Shuo Yu, Xiangjie Kong

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 185 pages

Journal ref Foundations and Trends in Signal Processing, Vol. 19, No. 4, pp 371-551. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04281 2025-11-07 cs.CV 57%

DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identification

Yujie Yang, Shuang Li, Jun Ye, Neng Dong, Fan Li, Huafeng Li

机构 * Kunming University of Science and Technology(昆明理工大学) Chongqing University of Post and Telecommunications(重庆邮电大学) China University of Mining Technology(中国矿业大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03332 2025-11-06 cs.CV 57%

Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge

Yi Yang, Yiming Xu, Timo Kaiser, Hao Cheng, Bodo Rosenhahn, Michael Ying Yang

机构 * Leibniz University Hannover(莱布尼茨汉诺威大学) University of Twente(特文特大学) University of Bath(巴斯大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18422 2025-11-06 cs.CV 57%

Breaking the Encoder Barrier for Seamless Video-Language Understanding

Handong Li, Yiyuan Zhang, Longteng Guo, Xiangyu Yue, Jing Liu

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MMLab, CUHK(CUHK MMLab) Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 23167-23176

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13174 2025-11-06 cs.CV 57%

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

Hao Cheng, Erjia Xiao, Yichi Wang, Chengyuan Yu, Mengshu Sun, Qiang Zhang, Jiahang Cao, Yijie Guo, Ning Liu, Kaidi Xu, Jize Zhang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) Beijing University of Technology(北京理工大学) Duke University(杜克大学) X-Humanoid Project(X-Humanoid 项目)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17664 2025-11-05 cs.CV cs.RO 57%

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(I2R, A*STAR) IPAL, CNRS IRL 2955, Singapore(IPAL, CNRS IRL 2955, 新加坡) CerCo, CNRS UMR 5549, Université Toulouse III(CerCo, CNRS UMR 5549, 法国图卢兹第三大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight; 43 pages, 17 figures, 16 tables; Project Page at https://talk2event.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01768 2025-11-04 cs.CV 57%

UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs

Zhe Liu, Jinghua Hou, Xiaoqing Ye, Jingdong Wang, Hengshuang Zhao, Xiang Bai

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology (HUST), Wuhan, China(华中科技大学电子信息与通信学院) School of Software Engineering, Huazhong University of Science and Technology (HUST), Wuhan, China(华中科技大学软件工程学院) Department of Computer Science, The University of Hong Kong, Hong Kong (HKU) SAR, China(香港大学计算机系) Baidu Inc., Beijing, China(百度公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00073 2025-11-04 cs.CV 57%

Habitat and Land Cover Change Detection in Alpine Protected Areas: A Comparison of AI Architectures

Harald Kristen, Daniel Kulmer, Manuela Hirschmugl

机构 * University of Graz(格拉茨大学) Joanneum Research(乔安姆研究)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15855 2025-11-04 q-bio.QM cs.AI cs.LG 57%

THFlow: A Temporally Hierarchical Flow Matching Framework for 3D Peptide Design

Dengdeng Huang, Shikui Tu

机构 * School of Computer Science Shanghai Jiao Tong University Shanghai, China(计算机科学学院 上海交通大学 上海中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27558 2025-11-03 cs.RO cs.AI cs.LG 57%

Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs

Sushil Samuel Dinesh, Shinkyu Park

机构 * Department of Electrical and Computer Engineering, King Abdullah University of Science and Technology (KAUST)(电气与计算机工程系,国王 Abdullah 科学与技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01930 2025-11-03 cs.CV 57%

PROFIT: A Specialized Optimizer for Deep Fine Tuning

Anirudh S Chakravarthy, Shuai Kyle Zheng, Xin Huang, Sachithra Hemachandra, Xiao Zhang, Yuning Chai, Zhao Chen

机构 * GM Cruise LLC

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments technical report, 23 pages, NeurIPS 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26027 2025-10-31 cs.CV 57%

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk, Mohsen Fayyaz

机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) L3S Research Center(L3S研究中心) Microsoft(微软公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01701 2025-10-31 cs.CV 57%

Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics

Naichuan Zheng, Yuchen Du, Hailun Xia, Zeyu Liang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22092 2025-10-29 cs.AI 57%

VIRAL: Vision-grounded Integration for Reward design And Learning

Valentin Cuzin-Rambaud, Emilien Komlenovic, Alexandre Faure, Bruno Yun

机构 * Université Claude Bernard Lyon 1(克莱尔伯恩大学里昂1分校) CNRS(国家科学研究中心) Ecole Centrale de Lyon(里昂中央理工学院) INSA Lyon(里昂工业高等学院) Université Lumière Lyon 2(里昂2大学卢米埃尔分校) LIRIS(图像研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23569 2025-10-28 cs.CV 57%

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He, Guo Chen, Fei Wu, Yu Qiao, Jiangmiao Pang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) The University of Tokyo(东京大学) Fudan University(复旦大学) Nanjing University(南京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23473 2025-10-28 cs.CV 57%

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang, Linxin Song, Runhao Fu, Hecheng Wang, Zongyuan Ge, Yuan Lu, Xuelian Cheng

机构 * Southeast University(东南大学) Monash University(墨尔本大学) Xiaohongshu Inc.(小红书公司) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23397 2025-10-28 cs.CV 57%

VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations

Lu Dong, Haiyu Zhang, Han Lin, Ziang Yan, Xiangyu Zeng, Hongjie Zhang, Yifei Huang, Yi Wang, Zhen-Hua Ling, Limin Wang, Yali Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23253 2025-10-28 cs.CV 57%

A Video Is Not Worth a Thousand Words

Sam Pollard, Michael Wray

机构 * University of Bristol(布里斯托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17924 2025-10-28 cs.CV 57%

Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation

Konstantin Egorov, Stepan Botman, Pavel Blinov, Galina Zubkova, Anton Ivaschenko, Alexander Kolsanov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) Samara State Medical University(萨马拉州医学大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与通信技术研究所可信人工智能研究中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACMMM 2025, Datasets track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20189 2025-10-27 cs.CV 57%

SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization

Xinyi Hu, Yuran Wang, Ruixu Zhang, Yue Li, Wenxuan Liu, Zheng Wang

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science(多媒体软件国家工程研究中心、人工智能研究院、计算机科学学院) Hubei Key Laboratory of Multimedia and Network Communication Engineering(多媒体与网络通信工程湖北省重点实验室) School of Mathematical Sciences, Peking University(北京大学数学科学学院) Tsinghua University(清华大学) School of Computer Science, Peking University(北京大学计算机科学学院) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08974 2025-10-27 cs.CV 57%

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

Elman Ghazaei, Erchan Aptoula

机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) Sabanci University(萨班奇大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18812 2025-10-27 cs.CV 57%

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21107 2025-10-27 cs.LG cs.AI cs.RO 57%

ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs

Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Gabor Karsai, Abhishek Dubey

机构 * Vanderbilt University(范德比大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Proceeding of the 39th Conference on Neural Information Processing Systems (NeurIPS'25). Code would be available at https://github.com/scope-lab-vu/ESCORT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20951 2025-10-27 cs.CV 57%

Generative Point Tracking with Flow Matching

Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis, Derek Nowrouzezahrai, Christopher Pal

机构 * McGill University(麦吉尔大学) Mila Stanford University(斯坦福大学) York University(约克大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://mtesfaldet.net/genpt_projpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20699 2025-10-24 q-fin.CP cs.AI 57%

Fusing Narrative Semantics for Financial Volatility Forecasting

Yaxuan Kong, Yoontae Hwang, Marcus Kaiser, Chris Vryonides, Roel Oomen, Stefan Zohren

机构 * University of Oxford(牛津大学) Pusan National University(釜山国立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments The 6th ACM International Conference on AI in Finance (ICAIF 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19574 2025-10-23 cs.CV cs.CR 57%

Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection

Ariana Yi, Ce Zhou, Liyang Xiao, Qiben Yan

机构 * Mission San Jose High School(Mission San Jose 高中) Missouri University of Science and Technology(密苏里科学与技术大学) Michigan State University(密歇根州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19560 2025-10-23 cs.CV 57%

HAD: Hierarchical Asymmetric Distillation to Bridge Spatio-Temporal Gaps in Event-Based Object Tracking

Yao Deng, Xian Zhong, Wenxuan Liu, Zhaofei Yu, Jingling Yuan, Tiejun Huang

机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology(武汉理工大学三亚科学教育创新园) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology(湖北省交通运输物联网重点实验室,计算机科学与人工智能学院,武汉理工大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21776 2025-10-23 cs.CV 57%

Video-R1: Reinforcing Video Reasoning in MLLMs

Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) CUHK (SZ)(香港中文大学(深圳)) Tsinghua University(清华大学) UCAS(中国科学院大学) CUHK HCCL(香港中文大学高性能计算实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025, Project page: https://github.com/tulerfeng/Video-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18726 2025-10-22 cs.CV 57%

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei, Yiwen He, Runzhe Wen, Chenxi Liao, Chengkang Jiang, An Ping, Shuo Gao, Suhan Wang, Zhaozhou Bian, Zijun Zhou, Jingyi Xie, Jiayi Zhou, Jing Wang, Yifan Yao, Weihao Xie, Yingshui Tan, Yanghai Wang, Qianqian Xie, Zhaoxiang Zhang, Jiaheng Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments https://github.com/NJU-LINK/IF-VidCap

详情

展开后加载摘要…

URL PDF HTML 收藏