arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-24 至 2025-09-24 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 13 篇

2504.08727 2025-09-24 cs.CV cs.AI cs.CY 84%

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

Boyang Deng, Songyou Peng, Kyle Genova, Gordon Wetzstein, Noah Snavely, Leonidas Guibas, Thomas Funkhouser

机构 * Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025, Project page: https://boyangdeng.com/visual-chronicles , second and third listed authors have equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13707 2025-09-24 cs.CV cs.AI 84%

EventVL: Understand Event Streams via Multimodal Large Language Model

Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) KU Leuven(根特大学) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) Carleton University(卡尔顿大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02889 2025-09-24 cs.CL cs.AI cs.CV cs.MM 83%

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

Xidong Wang, Dingjie Song, Shunian Chen, Junyin Chen, Zhenyang Cai, Chen Zhang, Lichao Sun, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Lehigh University(莱斯利大学) Meituan(美团)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04130 2025-09-24 cs.CV 79%

STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Jindong Jiang, Xiuyu Li, Zhijian Liu, Muyang Li, Guo Chen, Zhiqi Li, De-An Huang, Guilin Liu, Zhiding Yu, Kurt Keutzer, Sungjin Ahn, Jan Kautz, Hongxu Yin, Yao Lu, Song Han, Wonmin Byeon

机构 * NVIDIA(英伟达) Rutgers University(罗格斯大学) UC Berkeley(加州大学伯克利分校) MIT(麻省理工学院) Nanjing University(南京大学) KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18457 2025-09-24 cs.LG 78%

GluMind: Multimodal Parallel Attention and Knowledge Retention for Robust Cross-Population Blood Glucose Forecasting

Ebrahim Farahmand, Reza Rahimi Azghan, Nooshin Taheri Chatrudi, Velarie Yaa Ansu-Baidoo, Eric Kim, Gautham Krishna Gudur, Mohit Malu, Owen Krueger, Edison Thomaz, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh

机构 * Arizona State University(亚利桑那州立大学) University of Miami(迈阿密大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20201 2025-09-24 cs.CV cs.AI 73%

Injecting Explainability and Lightweight Design into Weakly Supervised Video Anomaly Detection Systems

Wen-Dong Jiang, Chih-Yung Chang, Hsiang-Chuan Chang, Ji-Yuan Chen, Diptendu Sinha Roy

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18230 2025-09-24 cs.AI cs.LG 70%

Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces

Zihan Dong, Xinyu Fan, Zixiang Tang, Yunqing Li

机构 * The University of Tokyo(东京大学) Lenovo US(联想美国) Advanced AI Technology Center(高级人工智能技术中心)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15173 2025-09-24 cs.CV cs.AI 62%

AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection

Zhipei Xu, Xuanyu Zhang, Qing Huang, Xing Zhou, Jian Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18187 2025-09-24 cs.CV cs.AI 62%

V-SenseDrive: A Privacy-Preserving Road Video and In-Vehicle Sensor Fusion Framework for Road Safety & Driver Behaviour Modelling

Muhammad Naveed, Nazia Perwaiz, Sidra Sultana, Mohaira Ahmad, Muhammad Moazam Fraz

机构 * School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST)(电气工程与计算机科学学院(SEECS),国立科学与技术大学(NUST))

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18183 2025-09-24 cs.CV cs.AI 62%

VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation

Jinyue Bian, Zhaoxing Zhang, Zhengyu Liang, Shiwei Zheng, Shengtao Zhang, Rong Shen, Chen Yang, Anzhou Hou

机构 * China, Beijing, Li Auto Inc.(中国北京李自动公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19245 2025-09-24 cs.CV 57%

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang, Elisa Ricci, Paolo Rota

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07560 2025-09-24 cs.RO cs.AI 57%

Socially Pertinent Robots in Gerontological Healthcare

Xavier Alameda-Pineda, Angus Addlesee, Daniel Hernández García, Chris Reinke, Soraya Arias, Federica Arrigoni, Alex Auternaud, Lauriane Blavette, Cigdem Beyan, Luis Gomez Camara, Ohad Cohen, Alessandro Conti, Sébastien Dacunha, Christian Dondrup, Yoav Ellinson, Francesco Ferro, Sharon Gannot, Florian Gras, Nancie Gunson, Radu Horaud, Moreno D'Incà, Imad Kimouche, Séverin Lemaignan, Oliver Lemon, Cyril Liotard, Luca Marchionni, Mordehay Moradi, Tomas Pajdla, Maribel Pino, Michal Polic, Matthieu Py, Ariel Rado, Bin Ren, Elisa Ricci, Anne-Sophie Rigaud, Paolo Rota, Marta Romeo, Nicu Sebe, Weronika Sieińska, Pinchas Tandeitnik, Francesco Tonini, Nicolas Turro, Timothée Wintz, Yanchao Yu

机构 * ERM Automatismes(ERM 自动化公司) PAL Robotics(PAL 机器人公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19169 2025-09-24 cs.RO 50%

MagiClaw: A Dual-Use, Vision-Based Soft Gripper for Bridging the Human Demonstration to Robotic Deployment Gap

Tianyu Wu, Xudong Han, Haoran Sun, Zishang Zhang, Bangchao Huang, Chaoyang Song, Fang Wan

机构 * Design + Learning Research Group(设计+学习研究组) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments 8 pages, 4 figures, accepted to Data@CoRL2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏