arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2209.15501 2022-10-11 cs.CV 57%

A Closer Look at Temporal Ordering in the Segmentation of Instructional Videos

Anil Batra, Shreyank N Gowda, Frank Keller, Laura Sevilla-Lara

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at BMVC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02953 2022-10-07 cs.CV 57%

Video Referring Expression Comprehension via Transformer with Content-aware Query

Ji Jiang, Meng Cao, Tengtao Song, Yuexian Zou

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13853 2022-09-29 cs.CV 57%

Thinking Hallucination for Video Captioning

Nasib Ullah, Partha Pratim Mohanta

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13583 2022-09-28 cs.CV 57%

Learning State-Aware Visual Representations from Audible Interactions

Himangi Mittal, Pedro Morgado, Unnat Jain, Abhinav Gupta

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments NeurIPS 2022. Code available at https://github.com/HimangiM/RepLAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12690 2022-09-28 cs.CV 57%

On Modality Bias Recognition and Reduction

Yangyang Guo, Liqiang Nie, Harry Cheng, Zhiyong Cheng, Mohan Kankanhalli, Alberto Del Bimbo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ToMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11990 2022-09-27 cs.CV 57%

Deep Neural Networks for Visual Reasoning

Thao Minh Le

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.08759 2022-09-20 cs.CV cs.IR 57%

Tree-based Text-Vision BERT for Video Search in Baidu Video Advertising

Tan Yu, Jie Liu, Yi Yang, Yi Li, Hongliang Fei, Ping Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments This revision is based on a manuscript submitted in October 2020, to ICDE 2021. We thank the Program Committee for their valuable comments

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05745 2022-09-14 cs.CL 57%

A virtual reality-based method for examining audiovisual prosody perception

Hartmut Meister, Isa Samira Winter, Moritz Waeachtler, Pascale Sandmann, Khaled Abdellatif

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00383 2022-09-07 cs.CV cs.IR 57%

ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

Naiyuan Liu, Xiaohan Wang, Xiaobo Li, Yi Yang, Yueting Zhuang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 1st Place in Ego4D Natural Language Queries Challenge, code is at https://github.com/NNNNAI/Ego4d_NLQ_2022_1st_Place_Solution

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.00786 2022-08-23 cs.CV 57%

Actor and Action Modular Network for Text-based Video Segmentation

Jianhua Yang, Yan Huang, Kai Niu, Linjiang Huang, Zhanyu Ma, Liang Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted By IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10400 2022-08-18 cs.CV 57%

Correspondence Matters for Video Referring Expression Comprehension

Meng Cao, Ji Jiang, Long Chen, Yuexian Zou

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06662 2022-08-16 cs.CV 57%

Self-Contained Entity Discovery from Captioned Videos

Melika Ayoughi, Pascal Mettes, Paul Groth

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05818 2022-08-12 cs.MM 57%

HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object Grounding

Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Wenqiao Zhang, Jiaxu Miao, Shiliang Pu, Fei Wu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14698 2022-08-08 cs.CV 57%

Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding

Jiachang Hao, Haifeng Sun, Pengfei Ren, Jingyu Wang, Qi Qi, Jianxin Liao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02450 2022-08-05 cs.CV 57%

Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-Identification

Xinyu Lin, Jinxing Li, Zeyu Ma, Huafeng Li, Shuang Li, Kaixiong Xu, Guangming Lu, David Zhang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 20973-20982

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01897 2022-08-04 cs.CV 57%

Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition

Mei Chee Leong, Haosong Zhang, Hui Li Tan, Liyuan Li, Joo Hwee Lim

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments The Ninth Workshop on Fine-Grained Visual Categorization (FGVC9) @ CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07706 2022-08-02 cs.CV 57%

Misinformation Detection in Social Media Video Posts

Kehan Wang, David Chan, Seth Z. Zhao, John Canny, Avideh Zakhor

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments We discovered an error in our dataset construction where retweets were not properly filtered. This resulted in test data leakage in training data, and the results reported are affected

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10765 2022-07-28 cs.RO cs.AI cs.LG 57%

Transporters with Visual Foresight for Solving Unseen Rearrangement Tasks

Hongtao Wu, Jikai Ye, Xin Meng, Chris Paxton, Gregory Chirikjian

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments IEEE IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12622 2022-07-27 cs.CV 57%

Multi-Attention Network for Compressed Video Referring Object Segmentation

Weidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06988 2022-07-27 cs.CV 57%

Event-guided Deblurring of Unknown Exposure Time Videos

Taewoo Kim, Jeongmin Lee, Lin Wang, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted in ECCV2022(Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05342 2022-07-22 cs.CV 57%

Video Graph Transformer for Video Question Answering

Junbin Xiao, Pan Zhou, Tat-Seng Chua, Shuicheng Yan

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments ECCV'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.09956 2022-07-21 cs.CV eess.IV 57%

Telepresence Video Quality Assessment

Zhenqiang Ying, Deepti Ghadiyaram, Alan Bovik

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.08380 2022-07-19 cs.CV 57%

Visual Representations of Physiological Signals for Fake Video Detection

Kalin Stefanov, Bhawna Paliwal, Abhinav Dhall

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07895 2022-07-19 cs.CV 57%

JPerceiver: Joint Perception Network for Depth, Pose and Layout Estimation in Driving Scenes

Haimei Zhao, Jing Zhang, Sen Zhang, Dacheng Tao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06237 2022-07-12 cs.CV 57%

Knowledge Distillation for Multi-Target Domain Adaptation in Real-Time Person Re-Identification

Félix Remigereau, Djebril Mekhazni, Sajjad Abdoli, Le Thanh Nguyen-Meidine, Rafael M. O. Cruz, Eric Granger

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 4 pages, 2 figures, submitted to ICIP2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04206 2022-07-12 cs.CL 57%

A Study of Syntactic Multi-Modality in Non-Autoregressive Machine Translation

Kexun Zhang, Rui Wang, Xu Tan, Junliang Guo, Yi Ren, Tao Qin, Tie-Yan Liu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10337 2022-07-11 cs.CV 57%

Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions

Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, Baining Guo

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Journal ref published in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07624 2022-07-06 cs.CV 57%

Attention Mechanisms in Computer Vision: A Survey

Meng-Hao Guo, Tian-Xing Xu, Jiang-Jiang Liu, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R. Martin, Ming-Ming Cheng, Shi-Min Hu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 27 pages, 9 figures

Journal ref Computational Visual Media, 2022, Vol. 8, No. 3, 331-368

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00579 2022-07-04 cs.CV cs.LG 57%

Video + CLIP Baseline for Ego4D Long-term Action Anticipation

Srijan Das, Michael S. Ryoo

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Secured second position in the Ego4D Challenge for Long-Term Action Anticipation track at CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00177 2022-07-04 cs.CV eess.IV 57%

Deep Motion Network for Freehand 3D Ultrasound Reconstruction

Mingyuan Luo, Xin Yang, Hongzhang Wang, Liwei Du, Dong Ni

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Early accepted by MICCAI-2022

详情

展开后加载摘要…

URL PDF HTML 收藏