arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2506.13956 2025-06-18 cs.CL cs.AI cs.RO 62%

ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection

Shang-Chi Tsai, Seiya Kawano, Angel Garcia Contreras, Koichiro Yoshino, Yun-Nung Chen

机构 * National Taiwan University(国立台湾大学) RIKEN(日本研究机构)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments IWSDS 2024 Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13654 2025-06-17 cs.CV cs.AI 62%

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

Shulin Tian, Ruiqi Wang, Hongming Guo, Penghao Wu, Yuhao Dong, Xiuying Wang, Jingkang Yang, Hao Zhang, Hongyuan Zhu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) A*STAR, Singapore(新加坡A*STAR) Simon Fraser University(西蒙弗雷泽大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://egolife-ai.github.io/Ego-R1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08185 2025-06-17 cs.CV cs.AI 62%

Agentic Surgical AI: Surgeon Style Fingerprinting and Privacy Risk Quantification via Discrete Diffusion in a Vision-Language-Action Framework

Huixin Zhan, Jason H. Moore

机构 * Cedars-Sinai Medical Center(西达萨凡纳医学中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08493 2025-06-11 cs.CV cs.MM 62%

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization

Qilin Yin, Wei Lu, Xiangyang Luo, Xiaochun Cao

机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University(计算机科学与工程学院、信息科技关键实验室、广东信息安全技术重点实验室、中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) School of Cyber Science and Technology, Shenzhen Campus, Sun Yat-sen University(网络安全科学与技术学院、深圳校区、中山大学)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08003 2025-06-10 cs.CV cs.AI 62%

Audio-Sync Video Generation with Multi-Stream Temporal Control

Shuchen Weng, Haojie Zheng, Zheng Chang, Si Li, Boxin Shi, Xinlong Wang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Nat’l Key Lab of General AI, School of Intelligence Science and Technology, Peking University(国家通用人工智能实验室,北京大学智能科学与技术学院) Nat’l Eng. Research Ctr. of Visual Tech., School of Computer Science, Peking University(国家视觉技术工程研究中心,北京大学计算机学院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06355 2025-06-10 cs.CY cs.CE cs.CL cs.CV 62%

LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment

Lingyao Li, Dawei Li, Zhenhui Ou, Xiaoran Xu, Jingxiao Liu, Zihui Ma, Runlong Yu, Min Deng

机构 * University of South Florida(佛罗里达州立大学) Arizona State University(亚利桑那州立大学) Massachusetts Institute of Technology(麻省理工学院) New York University(纽约大学) University of Alabama(阿拉巴马大学) Texas Tech University(德克萨斯科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21991 2025-06-09 cs.CV cs.AI 62%

A Lightweight Dual-Branch System for Weakly-Supervised Video Anomaly Detection on Consumer Edge Devices

Wen-Dong Jiang, Chih-Yung Chang, Ssu-Chi Kuai, Diptendu Sinha Roy

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This manuscript has been submitted to IEEE TCE and is under consideration for publication, with potential copyright transfer in the future

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03179 2025-06-05 cs.CV cs.AI 62%

Vid-SME: Membership Inference Attacks against Large Video Understanding Models

Qi Li, Runpeng Yu, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19475 2025-06-04 cs.CV cs.AI cs.LG 62%

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson, Robert Graham, Yash Vadi, Danilo Bzdok, Sebastian Lapuschkin, Lee Sharkey, Blake Aaron Richards

机构 * Mila Quebec(蒙特利尔大学) McGill University(麦吉尔大学) Meta Université de Montréal(蒙特利尔大学) Imperial College London(伦敦帝国理工学院) Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫 Heinrich Hertz 研究所) Technological University Dublin(都柏林技术大学) Apollo Research(Apollo 研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 4 pages, 3 figures, 9 tables. Oral and Tutorial at the CVPR Mechanistic Interpretability for Vision (MIV) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00928 2025-06-03 cs.CV cs.CL 62%

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

Olga Loginova, Sofía Ortega Loguinova

机构 * University of Trento(特伦托大学) Maastricht University(马斯特里赫特大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19753 2025-06-03 cs.GR cs.AI cs.CV 62%

A Survey on Event-driven 3D Reconstruction: Development under Different Categories

Chuanzhi Xu, Haoxian Zhou, Haodong Chen, Vera Chung, Qiang Qu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments We have decided not to submit this article and plan to withdraw it from public display. The content of this article will be presented in a more comprehensive form in another work

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16372 2025-05-23 cs.CV cs.AI 62%

Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition

Feng Liu, Bingyu Nan, Xuezhong Qian, Xiaolan Fu

机构 * School of Psychology,Shanghai Jiao Tong University(上海交通大学心理学学院) Jiangnan University(江南大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15447 2025-05-22 cs.CV cs.AI 62%

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Ziqiang Xu, Qi Dai, Tian Xie, Yifan Yang, Kai Qiu, DongDong Chen, Zuxuan Wu, Chong Luo

机构 * Fudan University(复旦大学) Microsoft(微软公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17821 2025-05-21 cs.CV cs.CL 62%

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

Xinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13123 2025-05-20 cs.CV cs.AI cs.LG 62%

Just Dance with $π$! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection

Snehashis Majhi, Giacomo D'Amicantonio, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, Egor Bondarev, Francois Bremond

机构 * INRIA Côte d’Azur University(科西嘉-阿兹尔大学) Woven by Toyota(丰田编织公司) Toyota Motor Europe(丰田欧洲公司) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16096 2025-05-20 q-bio.NC cs.AI cs.CV 62%

BrainPrompt: Multi-Level Brain Prompt Enhancement for Neurological Condition Identification

Jiaxing Xu, Kai He, Yue Tang, Wei Li, Mengcheng Lan, Xia Dong, Yiping Ke, Mengling Feng

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院) Saw Swee Hock School of Public Health at National University of Singapore, Singapore(新加坡国立大学 Saw Swee Hock 公共卫生学院) Xi’an Jiaotong University, Xi’an, Shaanxi(西安交通大学) S-Lab, Nanyang Technological University, Singapore(南洋理工大学 S 实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Early accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01645 2025-05-14 cs.CV cs.AI 62%

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

机构 * ByteDance(字节跳动) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Great Bay University(大湾大学) Shenzhen University(深圳大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18653 2025-05-13 cs.CV cs.AI 62%

When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation

Yuli Zhou, Guolei Sun, Yawei Li, Guo-Sen Xie, Luca Benini, Ender Konukoglu

机构 * Computer Vision Laboratory, ETH Zürich(苏黎世联邦理工学院计算机视觉实验室) Integrated System Laboratory, ETH Zürich(苏黎世联邦理工学院集成系统实验室) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Technical report. Accepted by Visual Intelligence. Code is released at https://github.com/zhoustan/SAM2-VCOS

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15106 2025-05-07 cs.CV cs.AI cs.LG 62%

About Time: Advances, Challenges, and Outlooks of Action Understanding

Alexandros Stergiou, Ronald Poppe

机构 * University of Twente(特文特大学) Utrecht University(乌特雷赫大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at the International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02768 2025-05-07 cs.CV cs.AI 62%

Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Jin Chen, Kaijing Ma, Haojian Huang, Han Fang, Hao Sun, Mehdi Hosseinzadeh, Zhe Liu

机构 * School of Computer Science, Duy Tan University(计算机科学学院,杜益坦大学) School of Computer Sciences, Universiti Sains Malaysia(计算机科学学院,马来亚大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03173 2025-05-07 cs.CV cs.AI 62%

RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph

Sameer Malik, Moyuru Yamada, Ayush Singh, Dishank Aggarwal

机构 * Fujitsu Research of India Private Limited(日本富士通印度研究私有限公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04369 2025-05-06 cs.CV cs.MM cs.NE 62%

SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition

Xiao Wang, Yao Rong, Zongzhen Wu, Lin Zhu, Bo Jiang, Jin Tang, Yonghong Tian

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) Beijing Institute Of Technology(北京理工大学) Peng Cheng Laboratory(鹏城实验室) National Engineering Laboratory for Video Technology, School of Electronics Engineering and Computer Science, Peking University(视频技术国家工程实验室,北京大学电子工程与计算机科学学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by IEEE Transactions on Cognitive and Developmental Systems (TCDS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14693 2025-05-06 cs.CV cs.AI 62%

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Enxin Song, Wenhao Chai, Weili Xu, Jianwen Xie, Yuxuan Liu, Gaoang Wang

机构 * Zhejiang University(浙江大学) University of Washington(华盛顿大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Lambda, Inc.(Lambda公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Code, docs, and benchmark are all avaliable at https://enxinsong.com/Video-MMLU-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01425 2025-05-05 cs.GR cs.AI cs.CV cs.LG cs.RO 62%

GENMO: A GENeralist Model for Human MOtion

Jiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe, Jan Kautz, Umar Iqbal, Ye Yuan

机构 * NVIDIA

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://research.nvidia.com/labs/dair/genmo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00755 2025-05-05 cs.CV cs.AI 62%

P2P-Insole: Human Pose Estimation Using Foot Pressure Distribution and Motion Sensors

Atsuya Watanabe, Ratna Aisuwarya, Lei Jing

机构 * Department of Computer Science and Engineering, University of Aizu(计算机科学与工程系,大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18689 2025-04-29 cs.CV cs.AI cs.LG 62%

HierSum: A Global and Local Attention Mechanism for Video Summarization

Apoorva Beedu, Irfan Essa

机构 * Georgia Institute of Technology(佐治亚理工学院) Google DeepMind(谷歌DeepMind)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08988 2025-04-28 cs.SD cs.MM eess.AS 62%

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Gaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi, Yuxin Peng, Anton van den Hengel, Jian Yang, Qingming Huang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Macquarie University(麦觉里大学) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) University of Adelaide(阿德莱德大学)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.MM、eess.AS

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17447 2025-04-25 cs.CV cs.AI 62%

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

De-An Huang, Subhashree Radhakrishnan, Zhiding Yu, Jan Kautz

机构 * NVIDIA(英伟达)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00622 2025-04-24 cs.CV cs.AI 62%

Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Xingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen, Adam Kortylewski, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Freiburg(弗赖堡大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICLR 2025 accepted paper. Project url: https://xingruiwang.github.io/projects/DynSuperCLEVR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15270 2025-04-22 cs.CV cs.CL 62%

An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes

Ji Qi, Yuan Yao, Yushi Bai, Bin Xu, Juanzi Li, Zhiyuan Liu, Tat-Seng Chua

机构 * Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏