arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2410.02362 2025-10-13 cs.CV cs.AI 62%

A Comprehensive Survey of Mamba Architectures for Medical Image Analysis: Classification, Segmentation, Restoration and Beyond

Shubhi Bansal, Sreeharish A, Madhava Prasath J, Manikandan S, Sreekanth Madisetty, Mohammad Zia Ur Rehman, Chandravardhan Singh Raghaw, Gaurav Duggal, Nagendra Kumar

机构 * Indian Institute of Technology Indore, India R.M.D. Engineering College, Kavaraipettai, India Jio Platforms Limited, India Birla Institute of Technology \& Science Pilani, India

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09200 2025-10-13 cs.CV cs.AI cs.HC 62%

Towards Safer and Understandable Driver Intention Prediction

Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar

机构 * IIIT Hyderabad(海得拉巴印度理工学院) Politecnico di Torino(托里诺理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07550 2025-10-10 cs.CV cs.AI 62%

TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility

Saman Motamed, Minghao Chen, Luc Van Gool, Iro Laina

机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06235 2025-10-09 eess.IV cs.AI cs.CV q-bio.NC 62%

Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)

Robert Scholz, Kunal Bagga, Christine Ahrends, Carlo Alberto Barbano

机构 * Université Paris Cité(巴黎Cité大学) University of Oxford(牛津大学) University of Turin(都灵大学) Universität Leipzig(莱比锡大学) Max Planck School of Cognition(马克斯·普朗克认知科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06077 2025-10-08 cs.CV cs.AI 62%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06040 2025-10-08 cs.CV cs.AI 62%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04819 2025-10-07 cs.CV cs.CL 62%

Visual Representations inside the Language Model

Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin, Winson Han, Ranjay Krishna

机构 * University of Washington(华盛顿大学) University of California Los Angeles(加州大学洛杉矶分校) Allen Institute for AI(人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 62%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04909 2025-10-01 cs.CV cs.AI 62%

HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding

Yuxuan Cai, Jiangning Zhang, Zhenye Gan, Qingdong He, Xiaobin Hu, Junwei Zhu, Yabiao Wang, Chengjie Wang, Zhucun Xue, Chaoyou Fu, Xinwei He, Xiang Bai

机构 * Fantasyele

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09079 2025-09-29 cs.CV cs.AI 62%

VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks

Xinlong Chen, Yuanxing Zhang, Yushuo Guan, Weihong Lin, Zekun Wang, Bohan Zeng, Yang Shi, Sihan Yang, Qiang Liu, Pengfei Wan, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所模式识别实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Peking University(北京大学) Nanjing University(南京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20703 2025-09-26 cs.RO cs.AI cs.CV 62%

Joint Flow Trajectory Optimization For Feasible Robot Motion Generation from Video Demonstrations

Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi

机构 * College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学) School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16421 2025-09-26 cs.CV cs.AI 62%

AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

Aiden Chang, Celso De Melo, Stephanie M. Lukin

机构 * University of Southern California(南加州大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025, 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17844 2025-09-26 cs.CL cs.AI 62%

THCM-CAL: Temporal-Hierarchical Causal Modelling with Conformal Calibration for Clinical Risk Prediction

Xin Zhang, Qiyu Wei, Yingjie Zhu, Fanyi Wu, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15173 2025-09-24 cs.CV cs.AI 62%

AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection

Zhipei Xu, Xuanyu Zhang, Qing Huang, Xing Zhou, Jian Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18187 2025-09-24 cs.CV cs.AI 62%

V-SenseDrive: A Privacy-Preserving Road Video and In-Vehicle Sensor Fusion Framework for Road Safety & Driver Behaviour Modelling

Muhammad Naveed, Nazia Perwaiz, Sidra Sultana, Mohaira Ahmad, Muhammad Moazam Fraz

机构 * School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST)(电气工程与计算机科学学院(SEECS),国立科学与技术大学(NUST))

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18183 2025-09-24 cs.CV cs.AI 62%

VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation

Jinyue Bian, Zhaoxing Zhang, Zhengyu Liang, Shiwei Zheng, Shengtao Zhang, Rong Shen, Chen Yang, Anzhou Hou

机构 * China, Beijing, Li Auto Inc.(中国北京李自动公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17888 2025-09-23 cs.CV cs.AI 62%

Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training

Divya Mereddy, Marcos Quinones-Grueiro, Ashwin T S, Eduardo Davalos, Gautam Biswas, Kent Etherton, Tyler Davis, Katelyn Kay, Jill Lear, Benjamin Goldberg

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13722 2025-09-18 cs.CV cs.AI 62%

Mitigating Query Selection Bias in Referring Video Object Segmentation

Dingwei Zhang, Dong Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology(香港科学大学) Nanjing Forestry University(南京林业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12876 2025-09-17 cs.CL cs.MM 62%

Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents

Fuyu Xing, Zimu Wang, Wei Wang, Haiyang Zhang

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL、cs.MM

Comments Accepted at INLG 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11232 2025-09-16 cs.CV cs.AI 62%

MIS-LSTM: Multichannel Image-Sequence LSTM for Sleep Quality and Stress Prediction

Seongwan Park, Jieun Woo, Siheon Yang

机构 * Sungkyunkwan University(釜山大学) Yeungnam University(延世大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICTC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02074 2025-09-10 cs.CV cs.AI 62%

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction and Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09632 2025-09-09 cs.CV cs.AI 62%

Preacher: Paper-to-Video Agentic System

Jingwei Liu, Ling Yang, Hao Luo, Fan Wang, Hongyan Li, Mengdi Wang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) Hupan Lab(虎扑实验室) National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05298 2025-09-09 cs.HC cs.AI cs.MM 62%

Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression

Rui Xi, Xianghan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted to the Proceedings of the 2025 International Conference on Artificial Intelligence and Virtual Reality (AIVR 2025). \c{opyright} 2025 Springer. This is the author-accepted manuscript. Rui Xi and Xianghan Wang contributed equally to this work. The final version will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 62%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01383 2025-09-03 cs.CV cs.MM 62%

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

Long Zhang, Peipei Song, Jianfeng Dong, Kun Li, Xun Yang

机构 * University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发式智能感知与认知实验室) Zhejiang Gongshang University(浙江工商大学) ReLER, CCAI, Zhejiang University(ReLER,中国计算机学会,浙江大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00210 2025-09-03 cs.CV cs.AI 62%

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu, Qinhan Lv, Xiying Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17932 2025-08-26 cs.CV cs.AI 62%

See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops

Zixuan Dong, Baoyun Peng, Yufei Wang, Lin Liu, Xinxin Dong, Yunlong Cao, Xiaodong Wang

机构 * College of Computer, National University of Defense Technology(计算机学院,国防科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17160 2025-08-26 cs.CV cs.AI 62%

Beyond Play and Pause: Turning GPT-4o Spatial Weakness into a Strength for In-Depth Interactive Video Learning

Sajad Goudarzi, Samaneh Zamanifard

机构 * Clemson University(克莱姆森大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 62%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16291 2025-08-25 cs.CV cs.MM 62%

Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment

Fengshun Wang, Qiurui Wang, Peilin Zhao

机构 * Capital University of \ Education Shanghai Jiao Tong University Shanghai China Shanghai Jiao Tong University

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏