arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-13 至 2026-01-13 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 8 篇

2601.06573 2026-01-13 cs.AI cs.MM 81%

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

QMAVIS:利用大多模态模型融合实现长视频音频理解

Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak, Chan Zhi Yang Jeriel, Jun Yang Lee, Yaohao Li

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University, Singapore(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 QMAVIS通过融合大型多模态模型、大型语言模型和语音识别模型,实现了长视频音频理解,展示了在VideoMME数据集上38.75%的性能提升,并在其他数据集上表现出竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06843 2026-01-13 cs.CV cs.CL 81%

Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models

边看边说:解锁多模态大语言模型的实时视频理解能力

Junyan Lin, Junlong Tong, Hao Wu, Jialiang Zhang, Jinming Liu, Xin Jin, Xiaoyu Shen

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Shanghai Jiao Tong University(上海交通大学) Ocean University of China(中国海洋大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出并行流式框架,通过三种设计解决多模态大语言模型在实时视频理解中的位置连续性约束问题,实现边看边说的实时交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22393 2026-01-13 cs.CV cs.LG 79%

Gems: Group Emotion Profiling Through Multimodal Situational Understanding

GEMS: 通过多模态情境理解进行群体情绪分析

Anubhav Kataria, Surbhi Madan, Shreya Ghosh, Tom Gedeon, Abhinav Dhall

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 GEMS通过多模态情境理解实现群体情绪分析,预测个体、群体和事件层面的情绪,提供更细粒度和整体的分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04376 2026-01-13 cs.CV 70%

Combining Facial Videos and Biosignals for Stress Estimation During Driving

结合面部视频和生物信号用于驾驶过程中的压力估计

Paraskevi Valergaki, Vassilis C. Nicodemou, Iason Oikonomidis, Antonis Argyros, Anastasios Roussos

机构 * Computer Science Department, University of Crete(塞萨洛尼基大学计算机科学系) Institute of Computer Science (ICS), Foundation for Research & Technology – Hellas (FORTH)(希腊基础研究与技术机构计算机科学研究所)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出结合面部视频和生物信号的多模态压力估计框架,通过跨模态注意力融合显著提升性能,适用于驾驶场景的压力监测。

Comments Under submission to ICPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06257 2026-01-13 q-bio.NC cs.AI cs.CV 62%

Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis

Gamma2Patterns: 深度认知注意力区域识别与伽马-阿尔法模式分析

Sobhana Jahan, Saydul Akbar Murad, Nick Rahimi, Noorbakhsh Amiri Golilarz

机构 * 1Department of Computer Science, The University of Alabama, Tuscaloosa, AL, USA 2School of Computing Sciences \& Computer Engineering, University of Southern Mississippi, Hattiesburg, MS, USA Email

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Gamma2Patterns通过多模式分析揭示深度注意力的神经区域和振荡特征,为人工智能系统中的注意力机制提供神经生理学基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07366 2026-01-13 cs.CV 57%

HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression

HiVid-Narrator:基于场景优先的ASR锚定压缩的分层视频叙事生成

Haoxuan Li, Mengyan Li, Junjun Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 HiVid-Narrator通过分阶段构建和SPA-Compressor压缩技术,在减少输入标记的同时生成高质量的电子商务视频叙事。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07290 2026-01-13 cs.CV 57%

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

VideoLoom:一种用于联合空间-时间理解的视频大语言模型

Jiapeng Shi, Junke Wang, Zuyao You, Bo He, Zuxuan Wu

机构 * Fudan University(复旦大学) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 VideoLoom是一种用于联合空间-时间理解的视频大语言模型,通过LoomData-8.7k数据集和LoomBench基准测试,在多个视频理解任务中取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07558 2026-01-13 cs.RO 50%

FlyCo: Foundation Model-Empowered Drones for Autonomous 3D Structure Scanning in Open-World Environments

FlyCo:基于基础模型的无人机自主3D结构扫描系统

Chen Feng, Guiyong Zheng, Tengkai Zhuang, Yongqian Wu, Fangzhan He, Haojia Li, Juepeng Zheng, Shaojie Shen, Boyu Zhou

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Department of Mechanical and Energy Engineering, Southern University of Science and Technology(南方科技大学机械与能源工程系) Differential Robotics, Hangzhou, China(杭州差分机器人)

专题命中 视频多模态 :multi-modal(abstract)

AI总结 FlyCo通过整合基础模型实现无人机自主3D扫描,提升开放世界环境下的目标定位与预测效率。

Comments 34 pages, 24 figures, 9 tables. Video: https://www.youtube.com/playlist?list=PLqjZjnqsCyl40rw3y15Yzc7Mdo-z1y2j8

详情

展开后加载摘要…

URL PDF HTML 收藏