arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

1904.03692 2019-04-09 cs.CV 57%

Unsupervised Domain Adaptation for Multispectral Pedestrian Detection

Dayan Guan, Xing Luo, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, George Vosselman, Michael Ying Yang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.06181 2019-03-20 cs.CV 57%

Dual Encoding for Zero-Example Video Retrieval

Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, Xun Wang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2019. Code and data are available at https://github.com/danieljf24/dual_encoding

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.01489 2019-03-06 cs.CV 57%

M-VAD Names: a Dataset for Video Captioning with Naming

Stefano Pini, Marcella Cornia, Federico Bolelli, Lorenzo Baraldi, Rita Cucchiara

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Source Code: https://github.com/aimagelab/mvad-names-dataset - Video Demo: https://youtu.be/dOvtAXbOOH4

Journal ref Multimedia Tools and Applications (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.03462 2019-01-14 cs.CV 57%

Analyzing Periodicity and Saliency for Adult Video Detection

Yizhi Liu, Xiaoyan Gu, Lei Huang, Junlin Ouyang, Miao Liao, Liangran Wu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.01451 2019-01-08 cs.CV 57%

Early Prediction of Alzheimer's Disease Dementia Based on Baseline Hippocampal MRI and 1-Year Follow-Up Cognitive Measures Using Deep Recurrent Neural Networks

Hongming Li, Yong Fan

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ISBI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.00484 2019-01-03 cs.CV 57%

Action2Vec: A Crossmodal Embedding Approach to Action Learning

Meera Hahn, Andrew Silva, James M. Rehg

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.04429 2018-12-12 cs.CV 57%

Face-Focused Cross-Stream Network for Deception Detection in Videos

Mingyu Ding, An Zhao, Zhiwu Lu, Tao Xiang, Ji-Rong Wen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.07014 2018-11-06 cs.CV 57%

To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression

Yitian Yuan, Tao Mei, Wenwu Zhu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06334 2018-11-05 cs.CV cs.LG 57%

Auxiliary Tasks in Multi-task Learning

Lukas Liebel, Marco Körner

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments fixed minor typesetting issue

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.06269 2018-10-30 cs.CV 57%

Learning Effective RGB-D Representations for Scene Recognition

Xinhang Song, Shuqiang Jiang, Luis Herranz, Chengpeng Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at IEEE Transactions on Image Processing

Journal ref IEEE Transactions on Image Processing, vol. 28, no. 2, pp. 980-993, Feb. 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.07110 2018-10-30 cs.CV 57%

Modality Distillation with Multiple Stream Networks for Action Recognition

Nuno Garcia, Pietro Morerio, Vittorio Murino

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2018; Supp. material at p.16; code available

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.06191 2018-09-21 cs.CV 57%

Multi Modal Convolutional Neural Networks for Brain Tumor Segmentation

Mehmet Aygün, Yusuf Hüseyin Şahin, Gözde Ünal

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03064 2018-09-18 cs.CV 57%

Recurrent CNN for 3D Gaze Estimation using Appearance and Shape Cues

Cristina Palmero, Javier Selva, Mohammad Ali Bagheri, Sergio Escalera

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Proc. of British Machine Vision Conference (BMVC), BMVC 2018. Errata: in pg.5 the camera matrices of the transformation matrix W should be interchanged (correct version: W=C_n*M*(C_o)^-1)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.05535 2018-08-17 stat.ML cs.CL cs.LG 57%

Combining time-series and textual data for taxi demand prediction in event areas: a deep learning approach

Filipe Rodrigues, Ioulia Markou, Francisco Pereira

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL

Comments 20 pages, 6 figures

Journal ref Rodrigues, F., Markou, I., Pereira, F. Combining time-series and textual data for taxi demand prediction in event areas: a deep learning approach. In Information Fusion, Elsevier, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.02559 2018-08-09 cs.CV 57%

A Joint Sequence Fusion Model for Video Question Answering and Retrieval

Youngjae Yu, Jongseok Kim, Gunhee Kim

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments To appear in ECCV 2018. 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.01837 2018-08-07 cs.CV 57%

Improving Temporal Interpolation of Head and Body Pose using Gaussian Process Regression in a Matrix Completion Setting

Stephanie Tan, Hayley Hung

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00108 2018-07-31 cs.CV 57%

Graph Distillation for Action Detection with Privileged Modalities

Zelun Luo, Jun-Ting Hsieh, Lu Jiang, Juan Carlos Niebles, Li Fei-Fei

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.08381 2018-07-24 cs.CV 57%

Pedestrian Trajectory Prediction with Structured Memory Hierarchies

Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments To appear in ECML-PKDD 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.04391 2018-06-13 cs.CV 57%

Qiniu Submission to ActivityNet Challenge 2018

Xiaoteng Zhang, Yixin Bao, Feiyun Zhang, Kai Hu, Yicheng Wang, Liang Zhu, Qinzhu He, Yining Lin, Jie Shao, Yao Peng

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 4 pages, 3 figures, CVPR workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.12098 2018-05-31 cs.CV 57%

Context-aware Cascade Attention-based RNN for Video Emotion Recognition

Man-Chin Sun, Shih-Huan Hsu, Min-Chun Yang, Jen-Hsien Chien

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.02489 2018-05-31 cs.HC cs.LG cs.SD eess.AS 57%

Transformer for Emotion Recognition

Jean-Benoit Delbrouck

专题命中 视频多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.07235 2018-05-08 cs.CV 57%

Tracking in Aerial Hyperspectral Videos using Deep Kernelized Correlation Filters

Burak Uzkent, Aneesh Rangnekar, Matthew J. Hoffman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00731 2018-05-03 cs.CL 57%

Exploring Emoji Usage and Prediction Through a Temporal Variation Lens

Francesco Barbieri, Luis Marujo, Pradeep Karuturi, William Brendel, Horacio Saggion

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Emojis @ ICWSM 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.02920 2018-04-24 cs.LG cs.AI cs.RO 57%

Vision-Based Multi-Task Manipulation for Inexpensive Robots Using End-To-End Learning from Demonstration

Rouhollah Rahmatizadeh, Pooya Abolghasemi, Ladislau Bölöni, Sergey Levine

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.06773 2018-03-20 cs.LG cs.AI cs.RO stat.ML 57%

Composable Deep Reinforcement Learning for Robotic Manipulation

Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, Sergey Levine

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Videos: https://sites.google.com/view/composing-real-world-policies/

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.06905 2018-03-06 cs.CV 57%

Learnable pooling with Context Gating for video classification

Antoine Miech, Ivan Laptev, Josef Sivic

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Presented at Youtube 8M CVPR17 Workshop. Kaggle Winning model. Under review for TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.00491 2018-02-05 cs.CV 57%

A New Registration Approach for Dynamic Analysis of Calcium Signals in Organs

Peixian Liang, Jianxu Chen, Pavel A. Brodskiy, Qinfeng Wu, Yejia Zhang, Yizhe Zhang, Lin Yang, Jeremiah J. Zartman, Danny Z. Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at ISBI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.04443 2017-12-14 cs.SI cs.AI cs.LG 57%

Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks

Bo Wu, Wen-Huang Cheng, Yongdong Zhang, Qiushi Huang, Jintao Li, Tao Mei

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments accepted in IJCAI-17

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.08097 2017-12-12 cs.CV 57%

Integrating both Visual and Audio Cues for Enhanced Video Caption

Wangli Hao, Zhaoxiang Zhang, He Guan, Guibo Zhu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Have some problems need to be handled

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.09550 2017-11-28 cs.CV cs.LG 57%

Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification

Xiang Long, Chuang Gan, Gerard de Melo, Jiajun Wu, Xiao Liu, Shilei Wen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments The backbone of the winner solution at ActivityNet Kinetics Challenge 2017

详情

展开后加载摘要…

URL PDF HTML 收藏