arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2101.02530 2023-03-07 cs.CV cs.LG eess.SP stat.AP stat.ML 74%

MSED: a multi-modal sleep event detection model for clinical sleep analysis

Alexander Neergaard Olesen, Poul Jennum, Emmanuel Mignot, Helge B. D. Sorensen

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 10 pages, 4 figures. Accepted for publication in IEEE Transactions on Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13606 2023-02-01 cs.CV 74%

Multi-video Moment Ranking with Multimodal Clue

Danyang Hou, Liang Pang, Yanyan Lan, Huawei Shen, Xueqi Cheng

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 9 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12460 2022-10-25 cs.CL 74%

Collaborative Reasoning on Multi-Modal Semantic Graphs for Video-Grounded Dialogue Generation

Xueliang Zhao, Yuxuan Wang, Chongyang Tao, Chenshuo Wang, Dongyan Zhao

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

Comments To appear at EMNLP 2022 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08113 2022-10-18 cs.CV 74%

Instance Segmentation with Cross-Modal Consistency

Alex Zihao Zhu, Vincent Casser, Reza Mahjourian, Henrik Kretzschmar, Sören Pirk

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments 8 pages, 9 figures, 5 tables. Presented at IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04246 2022-08-09 cs.CV cs.LG 74%

Snowpack Estimation in Key Mountainous Water Basins from Openly-Available, Multimodal Data Sources

Malachy Moran, Kayla Woputz, Derrick Hee, Manuela Girotto, Paolo D'Odorico, Ritwik Gupta, Daniel Feldman, Puya Vahabi, Alberto Todeschini, Colorado J Reed

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments Accepted Oral Presentation at CVPR 2022 MultiEarth

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.13274 2022-07-15 cs.LG cs.AI 74%

Evaluating Multimodal Interactive Agents

Josh Abramson, Arun Ahuja, Federico Carnevale, Petko Georgiev, Alex Goldin, Alden Hung, Jessica Landon, Timothy Lillicrap, Alistair Muldal, Blake Richards, Adam Santoro, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07096 2022-05-17 cs.CV cs.RO 74%

Multi-modal curb detection and filtering

Sandipan Das, Navid Mahabadi, Saikat Chatterjee, Maurice Fallon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15829 2022-03-31 cs.CV 74%

An EEG-Based Multi-Modal Emotion Database with Both Posed and Authentic Facial Actions for Emotion Analysis

Xiaotian Li, Xiang Zhang, Huiyuan Yang, Wenna Duan, Weiying Dai, Lijun Yin

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref FG2021(long Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07086 2022-03-15 cs.CV 74%

MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization

Alexander Kunitsyn, Maksim Kalashnikov, Maksim Dzabraev, Andrei Ivaniuta

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02252 2022-01-25 cs.CV 74%

Discourse Parsing in Videos: A Multi-modal Appraoch

Arjun R. Akula, Song-Chun Zhu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Accepted in CVPR 2019 Workshop on Language and Vision (Oral Presentation)

Journal ref CVPR 2019 Workshop on Language and Vision (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12341 2021-11-25 cs.CV 74%

EvDistill: Asynchronous Events to End-task Learning via Bidirectional Reconstruction-guided Cross-modal Knowledge Distillation

Lin Wang, Yujeong Chae, Sung-Hoon Yoon, Tae-Kyun Kim, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments CVPR 2021 (updated references in this version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12083 2021-11-24 cs.RO cs.CV cs.LG 74%

VISTA 2.0: An Open, Data-driven Simulator for Multimodal Sensing and Policy Learning for Autonomous Vehicles

Alexander Amini, Tsun-Hsuan Wang, Igor Gilitschenski, Wilko Schwarting, Zhijian Liu, Song Han, Sertac Karaman, Daniela Rus

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments First two authors contributed equally. Code and project website is available here: https://vista.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10699 2021-11-09 cs.CV 74%

MDMMT: Multidomain Multimodal Transformer for Video Retrieval

Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov, Aleksandr Petiushko

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Journal ref CVPR Workshops 2021: 3354-3363

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10211 2021-10-28 cs.CV 74%

Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning

Mandela Patrick, Yuki M. Asano, Bernie Huang, Ishan Misra, Florian Metze, Joao Henriques, Andrea Vedaldi

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted to ICCV 2021. Code at https://github.com/facebookresearch/GDT

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13992 2021-10-28 cs.CV cs.LG 74%

Leveraging Local Temporal Information for Multimodal Scene Classification

Saurabh Sahu, Palash Goyal

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14633 2021-05-03 physics.soc-ph cs.AI cs.LG cs.SI 74%

Modelling Urban Dynamics with Multi-Modal Graph Convolutional Networks

Krittika D'Silva, Jordan Cambe, Anastasios Noulas, Cecilia Mascolo, Adam Waksman

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 10 pages, 4 figures. arXiv admin note: substantial text overlap with arXiv:2104.13981

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06838 2021-02-18 cs.RO cs.CV 74%

Unified Multi-Modal Landmark Tracking for Tightly Coupled Lidar-Visual-Inertial Odometry

David Wisth, Marco Camurri, Sandipan Das, Maurice Fallon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Video: https://youtu.be/MjXYAHurWe8

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14462 2020-12-18 cs.LG astro-ph.IM cs.CV eess.IV eess.SP 74%

Deep Probabilistic Imaging: Uncertainty Quantification and Multi-modal Solution Characterization for Computational Imaging

He Sun, Katherine L. Bouman

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments This paper has been accepted to AAAI 2021. Keywords: Computational Imaging, Normalizing Flow, Uncertainty Quantification, Interferometry, MRI

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.02568 2020-09-08 cs.CV 74%

Multimodal Memorability: Modeling Effects of Semantics and Decay on Video Memorability

Anelise Newman, Camilo Fosco, Vincent Casser, Allen Lee, Barry McNamara, Aude Oliva

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments European Conference on Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03715 2020-08-21 eess.SP cs.DC cs.MM 74%

A Modular Approach for Synchronized Wireless Multimodal Multisensor Data Acquisition in Highly Dynamic Social Settings

Chirag Raman, Stephanie Tan, Hayley Hung

专题命中 视频多模态 :multimodal(title);分类 cs.MM

Comments 9 pages, 8 figures, Proceedings of the 28th ACM International Conference on Multimedia (MM '20), October 12--16, 2020, Seattle, WA, USA. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07544 2020-04-17 cs.CV eess.IV 74%

Multimodal and multiview distillation for real-time player detection on a football field

Anthony Cioppa, Adrien Deliège, Noor Ul Huda, Rikke Gade, Marc Van Droogenbroeck, Thomas B. Moeslund

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments Accepted for the CVSports workshop of CVPR 2020 ; 8 pages + references

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13594 2020-03-31 cs.CV 74%

Speech2Action: Cross-modal Supervision for Action Recognition

Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted to CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08830 2020-01-27 cs.SD cs.LG eess.AS eess.SP 74%

Scattering Features for Multimodal Gait Recognition

Srđan Kitić, Gilles Puy, Patrick Pérez, Philippe Gilberton

专题命中 视频多模态 :multimodal(title);分类 eess.AS

Comments Published at IEEE GlobalSIP 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01886 2019-02-07 cs.AI 74%

Situational Grounding within Multimodal Simulations

James Pustejovsky, Nikhil Krishnaswamy

专题命中 视频多模态 :multimodal(title);分类 cs.AI

Comments AAAI-19 Workshop on Games and Simulations for Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.00303 2018-12-04 cs.CV 74%

Multi-modal Capsule Routing for Actor and Action Video Segmentation Conditioned on Natural Language Queries

Bruce McIntosh, Kevin Duarte, Yogesh S Rawat, Mubarak Shah

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07212 2018-10-18 cs.CV 74%

Cross-Modal and Hierarchical Modeling of Video and Text

Bowen Zhang, Hexiang Hu, Fei Sha

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.00599 2018-10-02 cs.CV 74%

Unsupervised Trajectory Segmentation and Promoting of Multi-Modal Surgical Demonstrations

Zhenzhou Shao, Hongfa Zhao, Jiexin Xie, Ying Qu, Yong Guan, Jindong Tan

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06774 2018-04-19 cs.AI cs.NE cs.RO 74%

Encoding Longer-term Contextual Multi-modal Information in a Predictive Coding Model

Junpei Zhong, Tetsuya Ogata, Angelo Cangelosi

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments Submitted to ICDL/EpiRob 2018 (8th Joint IEEE International Conference on Development and Learning and on Epigenetic Robotics )

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.10330 2017-10-31 cs.CV 74%

Multi-modal Aggregation for Video Classification

Chen Chen, Xiaowei Zhao, Yang Liu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.07728 2017-10-30 cs.SI cs.CL cs.CY physics.soc-ph 74%

A Computational Framework for Multi-Modal Social Action Identification

Jason Anastasopoulos, Jake Ryland Williams

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏