arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2206.03248 2022-06-28 cs.CY cs.CL cs.LG 57%

Rites de Passage: Elucidating Displacement to Emplacement of Refugees on Twitter

Aparup Khatua, Wolfgang Nejdl

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments This work has been accepted to appear at HT'22-33rd ACM Conference on Hypertext and Social Media

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10149 2022-06-06 cs.LG cs.AI cs.RO 57%

Continuous Control with Action Quantization from Demonstrations

Robert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin, Anton Raichuk, Matthieu Geist, Olivier Pietquin

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to ICML 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15609 2022-06-02 cs.CV cs.IR 57%

BiC-Net: Learning Efficient Spatio-Temporal Relation for Text-Video Retrieval

Ning Han, Jingjing Chen, Chuhao Shi, Yawen Zeng, Guangyi Xiao, Hao Chen

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10431 2022-05-30 cs.LG cs.AI 57%

Learning Dense Reward with Temporal Variant Self-Supervision

Yuning Wu, Jieliang Luo, Hui Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 4 pages, 6 figures, accepted to ICRA 2022 RL for Contact-Rich Manipulation Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08508 2022-05-18 cs.CV 57%

A CLIP-Hitchhiker's Guide to Long Video Retrieval

Max Bain, Arsha Nagrani, Gül Varol, Andrew Zisserman

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06530 2022-05-16 cs.CV 57%

Modeling Semantic Composition with Syntactic Hypergraph for Video Question Answering

Zenan Xu, Wanjun Zhong, Qinliang Su, Zijing Ou, Fuwei Zhang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 11pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05895 2022-05-13 cs.CV 57%

Weakly-Supervised Action Detection Guided by Audio Narration

Keren Ye, Adriana Kovashka

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments To appear, in Joint 1st Ego4D and 10th EPIC Workshop, held in conjunction with the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03297 2022-05-09 cs.IR cs.MM 57%

Implicit semantic-based personalized micro-videos recommendation

Bo Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03124 2022-05-09 cs.CV 57%

A High-Accuracy Unsupervised Person Re-identification Method Using Auxiliary Information Mined from Datasets

Hehan Teng, Tao He, Yuchen Guo, Guiguang Ding

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00823 2022-05-03 cs.CV cs.IR 57%

CenterCLIP: Token Clustering for Efficient Text-Video Retrieval

Shuai Zhao, Linchao Zhu, Xiaohan Wang, Yi Yang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments accepted by SIGIR 2022, code is at https://github.com/mzhaoshuai/CenterCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11410 2022-04-26 cs.CV 57%

Single Object Tracking Research: A Survey

Ruize Han, Wei Feng, Qing Guo, Qinghua Hu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 32 pages, in Chinese survey paper, Chinese Journal of Computers 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13241 2022-04-26 cs.CV 57%

Learning from Temporal Gradient for Semi-supervised Action Recognition

Junfei Xiao, Longlong Jing, Lin Zhang, Ju He, Qi She, Zongwei Zhou, Alan Yuille, Yingwei Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16784 2022-04-01 cs.CV 57%

Video-Text Representation Learning via Differentiable Weak Temporal Alignment

Dohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh, Kyoung-Woon On, Eun-Sol Kim, Hyunwoo J. Kim

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15969 2022-03-31 cs.CV 57%

Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation

Guang Feng, Lihe Zhang, Zhiwei Hu, Huchuan Lu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14825 2022-03-29 cs.CV 57%

HDR Reconstruction from Bracketed Exposures and Events

Richard Shaw, Sibi Catley-Chandar, Ales Leonardis, Eduardo Perez-Pellitero

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07303 2022-03-15 cs.CV 57%

All in One: Exploring Unified Video-Language Pre-training

Alex Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, Xiaohu Qie, Mike Zheng Shou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 18 pages. 11 figures. Code: https://github.com/showlab/all-in-one

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.00487 2022-03-15 cs.CV 57%

Language as Queries for Referring Video Object Segmentation

Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, Ping Luo

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 14 pages, accepted by CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03253 2022-03-08 cs.CV 57%

Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal Information

Lingfeng Yang, Xiang Li, Renjie Song, Borui Zhao, Juntian Tao, Shihao Zhou, Jiajun Liang, Jian Yang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in CVPR22

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12396 2022-02-21 cs.CV 57%

Using Motion History Images with 3D Convolutional Networks in Isolated Sign Language Recognition

Ozge Mercanoglu Sincan, Hacer Yalim Keles

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05102 2022-02-16 cs.CV cs.LG eess.IV 57%

Self-Supervised Multisensor Change Detection

Sudipan Saha, Patrick Ebel, Xiao Xiang Zhu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01176 2022-01-20 cs.CV 57%

Overcoming the Domain Gap in Neural Action Representations

Semih Günel, Florian Aymanns, Sina Honari, Pavan Ramdya, Pascal Fua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05651 2022-01-19 cs.AI cs.LG 57%

CLUE: Contextualised Unified Explainable Learning of User Engagement in Video Lectures

Sujit Roy, Gnaneswara Rao Gorle, Vishal Gaur, Haider Raza, Shoaib Jameel

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08822 2021-12-28 cs.CV 57%

Non-contact Pain Recognition from Video Sequences with Remote Physiological Measurements Prediction

Ruijing Yang, Ziyu Guan, Zitong Yu, Xiaoyi Feng, Jinye Peng, Guoying Zhao

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments IJCAI 2021

Journal ref https://www.ijcai.org/proceedings/2021/170

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01918 2021-12-15 cs.CV cs.IT math.IT 57%

Multi-mode Core Tensor Factorization based Low-Rankness and Its Applications to Tensor Completion

Haijin Zeng

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.08974 2021-12-10 cs.CV 57%

ISSAFE: Improving Semantic Segmentation in Accidents by Fusing Event-based Data

Jiaming Zhang, Kailun Yang, Rainer Stiefelhagen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01479 2021-12-07 cs.CV 57%

Learning Spatial-Temporal Graphs for Active Speaker Detection

Sourya Roy, Kyle Min, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10596 2021-12-03 cs.CV cs.LG 57%

Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

Reuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin, Bryan Russell

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2021 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14547 2021-12-01 cs.CV 57%

LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question Answering

Jingjing Jiang, Ziyi Liu, Nanning Zheng

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 11 pages, 5 figures, Code: https://github.com/jingjing12110/LiVLR-VideoQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14595 2021-11-30 cs.CV 57%

Overcoming the Domain Gap in Contrastive Learning of Neural Action Representations

Semih Günel, Florian Aymanns, Sina Honari, Pavan Ramdya, Pascal Fua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted into NeurIPS 2021 Workshop: Self-Supervised Learning - Theory and Practice

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13324 2021-11-29 cs.CV 57%

Hierarchical Motion Encoder-Decoder Network for Trajectory Forecasting

Qifan Xue, Shengyi Li, Xuanpeng Li, Jingwen Zhao, Weigong Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏