arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 15 篇

2509.04086 2025-10-28 cs.CV cs.MM 84%

TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph

Yaru Chen, Faegheh Sardari, Peiliang Zhang, Ruohao Guo, Yang Xiang, Zhenbo Li, Wenwu Wang

机构 * Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(视觉、语音和信号处理中心(CVSSP),萨里大学) School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学) National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,北京大学智能科学与技术学院) College of Information and Electrical Engineering, China Agricultural University(信息与电子工程学院,中国农业大学)

专题命中 视频多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 81%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21761 2025-10-28 cs.RO cs.AI cs.CV 81%

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino

机构 * Division of Information Science, NAIST(NAIST信息科学系) Guardian Robot Project, RIKEN(RIKEN守护机器人项目)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06499 2025-10-28 q-bio.NC cs.AI 79%

The ISLab Solution to the Algonauts Challenge 2025: A Multimodal Deep Learning Approach to Brain Response Prediction

Andrea Corsico, Giorgia Rigamonti, Simone Zini, Luigi Celona, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01481 2025-10-28 cs.CV cs.LG 74%

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09732 2025-10-28 cs.LG q-bio.PE stat.AP 71%

Continental-scale habitat distribution modelling with multimodal earth observation foundation models

Sara Si-Moussi, Stephan Hennekens, Sander Mucher, Stan Los, Yoann Cartier, Borja Jiménez-Alfaro, Fabio Attorre, Jens-Christian Svenning, Wilfried Thuiller

专题命中 视频多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21786 2025-10-28 cs.CV cs.AI cs.MM 67%

EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction

Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong

机构 * Beihang University(北京航空航天大学) University of Science and Technology Beijing(北京科技大学) School of Computer Science and Engineering(计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 15 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22602 2025-10-28 cs.CL cs.AI cs.CY 62%

Personal Care Utility (PCU): Building the Health Infrastructure for Everyday Insight and Guidance

Mahyar Abbasian, Ramesh Jain

机构 * University of California, Irvine(加州大学尔湾分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 2 figures, 1 table, Journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21813 2025-10-28 cs.CV cs.AI cs.LG 62%

SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling

Samuel J. Barrett, Docko Sow

机构 * LGND AI Tolbi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23569 2025-10-28 cs.CV 57%

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He, Guo Chen, Fei Wu, Yu Qiao, Jiangmiao Pang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) The University of Tokyo(东京大学) Fudan University(复旦大学) Nanjing University(南京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23473 2025-10-28 cs.CV 57%

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang, Linxin Song, Runhao Fu, Hecheng Wang, Zongyuan Ge, Yuan Lu, Xuelian Cheng

机构 * Southeast University(东南大学) Monash University(墨尔本大学) Xiaohongshu Inc.(小红书公司) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23397 2025-10-28 cs.CV 57%

VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations

Lu Dong, Haiyu Zhang, Han Lin, Ziang Yan, Xiangyu Zeng, Hongjie Zhang, Yifei Huang, Yi Wang, Zhen-Hua Ling, Limin Wang, Yali Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23253 2025-10-28 cs.CV 57%

A Video Is Not Worth a Thousand Words

Sam Pollard, Michael Wray

机构 * University of Bristol(布里斯托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17924 2025-10-28 cs.CV 57%

Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation

Konstantin Egorov, Stepan Botman, Pavel Blinov, Galina Zubkova, Anton Ivaschenko, Alexander Kolsanov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) Samara State Medical University(萨马拉州医学大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与通信技术研究所可信人工智能研究中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACMMM 2025, Datasets track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22339 2025-10-28 cs.RO 50%

Estimating Continuum Robot Shape under External Loading using Spatiotemporal Neural Networks

Enyi Wang, Zhen Deng, Chuanchuan Pan, Bingwei He, Jianwei Zhang

机构 * Hamlyn Centre for Robotic Surgery, Institute of Global Health Innovation, Imperial College London(帝国理工学院伦敦校区全球健康创新研究所机器人手术中心) Department of Mechanical Engineering and Automation, Fuzhou University(福州大学机械工程与自动化学院) TAMS Group, Informatics, University of Hamburg(汉堡大学信息学院TAMS集团)

专题命中 视频多模态 :multi-modal(abstract)

Comments 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏