arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2006.09979 2020-06-18 cs.IR cs.LG 78%

I know why you like this movie: Interpretable Efficient Multimodal Recommender

Barbara Rychalska, Dominika Basaj, Jacek Dąbrowski, Michał Daniluk

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00028 2020-06-02 cs.RO 78%

Multi-modal Transfer Learning for Grasping Transparent and Specular Objects

Thomas Weng, Amith Pallankize, Yimin Tang, Oliver Kroemer, David Held

专题命中 视频多模态 :multi-modal(title,abstract)

Comments RA-L with presentation at ICRA 2020

Journal ref IEEE ROBOTICS AND AUTOMATION LETTERS, VOL. 5, NO. 3, JULY 2020. 3791-3798

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.01023 2020-04-03 cs.MM cs.CV cs.CY cs.SD eess.AS 78%

Multi-Modal Video Forensic Platform for Investigating Post-Terrorist Attack Scenarios

Alexander Schindler, Andrew Lindley, Anahid Jalali, Martin Boyer, Sergiu Gordea, Ross King

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.MM、eess.AS

Journal ref In Proceedings of the 11th ACM Multimedia Systems Conference (MMSys2020), June 06-11, 2020, Istanbul, Turkey

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.10643 2020-03-25 cs.NI cs.SI stat.AP stat.ML 78%

DeepSIP: A System for Predicting Service Impact of Network Failure by Temporal Multimodal CNN

Yoichi Matsuo, Tatsuaki Kimura, Ken Nishimatsu

专题命中 视频多模态 :multimodal(title,abstract)

Comments to appear in IEEE/IFIP International Workshop on Analytics for Network and Service Management (AnNet 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.02632 2020-03-06 eess.SP 78%

STAR: Spatio-Temporal Prediction of Air Quality Using A Multimodal Approach

Tien-Cuong Bui, Joonyoung Kim, Taewoo Kang, Donghyeon Lee, Junyoung Choi, Insoon Yang, Kyomin Jung, Sang Kyun Cha

专题命中 视频多模态 :multimodal(title,abstract)

Comments 18 pages, 9 figures, Intelligent System Conference (Intellisys 2020 - Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.00444 2020-01-20 cs.HC 78%

On Assessing Driver Awareness of Situational Criticalities: Multi-modal Bio-sensing and Vision-based Analysis, Evaluations, and Insights

Siddharth Siddharth, Mohan M. Trivedi

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Journal article published in MDPI Brain Sciences. arXiv admin note: text overlap with arXiv:1905.00503

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05629 2019-08-28 cs.CY 78%

A blockchain-based user-centric emission monitoring and trading system for multi-modal mobility

Johannes Eckert, David López, Carlos Lima Azevedo, Bilal Farooq

专题命中 视频多模态 :multi-modal(title,abstract)

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.07012 2019-07-09 cs.RO 78%

Understanding of Object Manipulation Actions Using Human Multi-Modal Sensory Data

Bahareh Abbasi, Ehsan Noohi, Sina Parastegari, Milos Zefran

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11688 2019-06-10 q-bio.NC 78%

Disrupted core-periphery structure of multimodal brain networks in Alzheimer's Disease

Jeremy Guillon, Mario Chavez, Federico Battiston, Yohan Attal, Valentina La Corte, Michel Thiebaut de Schotten, Bruno Dubois, Denis Schwartz, Olivier Colliot, Fabrizio De Vico Fallani

专题命中 视频多模态 :multimodal(title,abstract)

Comments 5 figures, 1 table, 1 supplementary figure, 2 supplementary tables

Journal ref Network Neuroscience 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.07039 2019-05-20 cs.LG cs.HC eess.SP stat.ML 78%

Utilizing Deep Learning Towards Multi-modal Bio-sensing and Vision-based Affective Computing

Siddharth Siddharth, Tzyy-Ping Jung, Terrence J. Sejnowski

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted for publication in IEEE Transactions on Affective Computing. This version on the arXiv is the updated version of the same manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07319 2019-04-17 cs.RO cs.LG 78%

Learning Probabilistic Multi-Modal Actor Models for Vision-Based Robotic Grasping

Mengyuan Yan, Adrian Li, Mrinal Kalakrishnan, Peter Pastor

专题命中 视频多模态 :multi-modal(title,abstract)

Journal ref The 2019 International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.02570 2019-04-05 eess.SP 78%

Can multimodal sensing detect and localize transient events?

Kasthuri Jayarajah, Vigneshwaran Subbaraju, Noel Athaide, Lakmal Meegahapola, Andrew Tan, Archan Misra

专题命中 视频多模态 :multimodal(title,abstract)

Journal ref SPIE Defense+Security 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.05566 2019-02-18 cs.IR 78%

Interest-Related Item Similarity Model Based on Multimodal Data for Top-N Recommendation

Junmei Lv, Bin Song, Jie Guo, Xiaojiang Du, Mohsen Guizani

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01117 2019-02-05 cs.HC cs.RO 78%

Exploring Temporal Dependencies in Multimodal Referring Expressions with Mixed Reality

Elena Sibirtseva, Ali Ghadirzadeh, Iolanda Leite, Mårten Björkman, Danica Kragic

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08093 2019-01-25 q-bio.NC 78%

Decoding multimodal behavior using time differences of MEG events

Ohad Felsenstein, Idan Tal, Michal Ben-Shachar, Moshe Abeles, Gal Chechik

专题命中 视频多模态 :multimodal(title,abstract)

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.10027 2018-12-03 cs.CY 78%

Multimodal Classification of Stressful Environments in Visually Impaired Mobility Using EEG and Peripheral Biosignals

Charalampos Saitis, Kyriaki Kalimeri

专题命中 视频多模态 :multimodal(title,abstract)

Comments IEEE Transactions on Affective Computing 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.09452 2018-06-22 cs.HC 78%

Multi-modal Approach for Affective Computing

Siddharth Siddharth, Tzyy-Ping Jung, Terrence J. Sejnowski

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Published in IEEE 40th International Engineering in Medicine and Biology Conference (EMBC) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.08947 2018-03-28 stat.AP cs.IT math.IT 78%

Sequential Event Detection Using Multimodal Data in Nonstationary Environments

Taposh Banerjee, Gene Whipps, Prudhvi Gurram, Vahid Tarokh

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.09232 2018-01-30 physics.med-ph stat.AP 78%

Multimodal Functional and Structural Brain Connectivity Analysis in Autism: A Preliminary Integrated Approach with EEG, fMRI and DTI

Bogdan Alexandru Cociu, Saptarshi Das, Lucia Billeci, Wasifa Jamal, Koushik Maharatna, Sara Calderoni, Antonio Narzisi, Filippo Muratori

专题命中 视频多模态 :multimodal(title,abstract)

Comments 14 pages, 14 figures, IEEE Transactions on Cognitive and Developmental Systems, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.10479 2017-11-27 cs.RO cs.LG 78%

Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets

Karol Hausman, Yevgen Chebotar, Stefan Schaal, Gaurav Sukhatme, Joseph Lim

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Paper accepted to NIPS 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.06039 2017-09-19 cs.RO 78%

Why did the Robot Cross the Road? - Learning from Multi-Modal Sensor Data for Autonomous Road Crossing

Noha Radwan, Wera Winterhalter, Christian Dornhege, Wolfram Burgard

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Video: https://www.youtube.com/watch?v=N1IhHHkUzYg Dataset: http://www2.informatik.uni-freiburg.de/~radwann/freiburg_street_crossing_dataset.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.05616 2017-03-17 cs.HC 78%

Multimodal Language Specification for Human Adaptive Mechatronics

Fernando Ferri, Arianna D'Ulizia, Patrizia Grifoni

专题命中 视频多模态 :multimodal(title,abstract)

Comments 11 pages, 4 figures

Journal ref Journal of Next Generation Information Technology (JNIT), Volume 3, Number 1, November 2012

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08492 2016-11-28 cs.HC 78%

A Multimodal Approach to Estimating Vigilance Using EEG and Forehead EOG

Wei-Long Zheng, Bao-Liang Lu

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.01209 2016-10-06 cs.CY 78%

Towards Air Quality Estimation Using Collected Multimodal Environmental Data

Anastasia Moumtzidou, Symeon Papadopoulos, Stefanos Vrochidis, Ioannis Kompatsiaris, Konstantinos Kourtidis, George Hloupis, Ilias Stavrakas, Konstantina Papachristopoulou, Christodoulos Keratidis

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.00306 2016-05-17 stat.AP cs.SI 78%

Multimodal Event Detection in Twitter Hashtag Networks

Yasin Yilmaz, Alfred Hero

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.4620 2016-02-26 q-bio.QM 78%

Integrated multimodal network approach to PET and MRI based on multidimensional persistent homology

Hyekyoung Lee, Hyejin Kang, Moo K. Chung, Seonhee Lim, Bung-Nyun Kim, Dong Soo Lee

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.04364 2016-02-16 cs.LG 78%

Look, Listen and Learn - A Multimodal LSTM for Speaker Identification

Jimmy Ren, Yongtao Hu, Yu-Wing Tai, Chuan Wang, Li Xu, Wenxiu Sun, Qiong Yan

专题命中 视频多模态 :multimodal(title,abstract)

Comments The 30th AAAI Conference on Artificial Intelligence (AAAI-16)

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.1274 2014-11-06 physics.soc-ph 78%

Anatomy and efficiency of urban multimodal mobility

Riccardo Gallotti, Marc Barthelemy

专题命中 视频多模态 :multimodal(title,abstract)

Comments 9 pages, 7 figures +supplementary information (9 pages and 9 figures)

Journal ref Scientific Reports 4:6911 (2014)

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.6425 2013-11-26 math.OC cs.LG stat.ML 78%

Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching

Marcelo Fiori, Pablo Sprechmann, Joshua Vogelstein, Pablo Musé, Guillermo Sapiro

专题命中 视频多模态 :multimodal(title,abstract)

Comments NIPS 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19075 2026-05-20 cs.CV cs.AI 77%

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

CRAFT: 基于批评的自适应关键帧目标定位用于多模态视频问答

Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi, Pengyu Yan, Akhil Gorugantu, David Doermann

机构 * University at Buffalo(布法罗大学) New York University(纽约大学)

专题命中 视频多模态 :multimodal(title,comments);分类 cs.CV、cs.AI

AI总结 该研究提出CRAFT方法,通过动态关键帧选择、每视频ASR与多语言回退以及混合批评循环,迭代验证和修复声明,最终实现多模态视频问答的准确证据聚合。

Comments Accepted at ACL 2026 Multimodal Augmented Generation via MultimodAl Retrieval Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏