arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2106.08408 2021-06-17 cs.CV eess.IV 57%

Seeing Through Clouds in Satellite Images

Mingmin Zhao, Peder A. Olsen, Ranveer Chandra

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07166 2021-06-15 cs.CV 57%

2rd Place Solutions in the HC-STVG track of Person in Context Challenge 2021

YiYu, XinyingWang, WeiHu, XunLuo, ChengLi

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06561 2021-06-15 cs.CV cs.GR cs.LG 57%

GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation (works for videos too!)

Min Jin Chong, David Forsyth

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments code is here https://github.com/mchong6/GANsNRoses

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06138 2021-06-14 cs.CV 57%

Team RUC_AIM3 Technical Report at ActivityNet 2021: Entities Object Localization

Ludan Ruan, Jieting Chen, Yuqing Song, Shizhe Chen, Qin Jin

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.10110 2021-06-08 cs.CV 57%

Guidance and Teaching Network for Video Salient Object Detection

Yingxia Jiao, Xiao Wang, Yu-Cheng Chou, Shouyuan Yang, Ge-Peng Ji, Rong Zhu, Ge Gao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at IEEE ICIP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13533 2021-05-31 cs.CV cs.HC cs.LG eess.SP 57%

Inertial Sensor Data To Image Encoding For Human Action Recognition

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07667 2021-05-18 cs.CV 57%

AudioVisual Video Summarization

Bin Zhao, Maoguo Gong, Xuelong Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06754 2021-05-17 cs.CV 57%

Learning Group Activities from Skeletons without Individual Action Labels

Fabio Zappardino, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05226 2021-05-12 cs.CV 57%

Home Action Genome: Cooperative Compositional Action Understanding

Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, Juan Carlos Niebles

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments CVPR '21

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05066 2021-05-12 cs.CV 57%

ChaLearn LAP Large Scale Signer Independent Isolated Sign Language Recognition Challenge: Design, Results and Future Research

Ozge Mercanoglu Sincan, Julio C. S. Jacques Junior, Sergio Escalera, Hacer Yalim Keles

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Preprint of the accepted paper at ChaLearn Looking at People Sign Language Recognition in the Wild Workshop at CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.01553 2021-05-05 cs.CV 57%

Combining Supervised and Un-supervised Learning for Automatic Citrus Segmentation

Heqing Huang, Tongbin Huang, Zhen Li, Zhiwei Wei, Shilei Lv

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 7 pages,4 figures,Prepare for submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02391 2021-04-07 cs.CV 57%

Weakly Supervised Video Salient Object Detection

Wangbo Zhao, Jing Zhang, Long Li, Nick Barnes, Nian Liu, Junwei Han

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Journal ref 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05847 2021-04-01 cs.CV 57%

ZeroScatter: Domain Transfer for Long Distance Imaging and Vision through Scattering Media

Zheng Shi, Ethan Tseng, Mario Bijelic, Werner Ritter, Felix Heide

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 2021 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), project page available at https://light.princeton.edu/publication/zeroscatter/

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15975 2021-03-31 cs.AI 57%

Platform for Situated Intelligence

Dan Bohus, Sean Andrist, Ashley Feniello, Nick Saw, Mihai Jalobeanu, Patrick Sweeney, Anne Loomis Thompson, Eric Horvitz

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 29 pages, 14 figures, Microsoft Research Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02604 2021-03-29 cs.RO cs.AI cs.LG 57%

Leveraging Post Hoc Context for Faster Learning in Bandit Settings with Applications in Robot-Assisted Feeding

Ethan K. Gordon, Sumegh Roychowdhury, Tapomayukh Bhattacharjee, Kevin Jamieson, Siddhartha S. Srinivasa

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 6 pages + references, 5 figures, to appear in ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11555 2021-03-23 cs.CV 57%

Context-aware Biaffine Localizing Network for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, Yulai Xie

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05488 2021-02-23 cs.CV 57%

EEV: A Large-Scale Dataset for Studying Evoked Expressions from Video

Jennifer J. Sun, Ting Liu, Alan S. Cowen, Florian Schroff, Hartwig Adam, Gautam Prasad

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Data subset at https://github.com/google-research-datasets/eev

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.00027 2021-02-02 cs.CV cs.RO 57%

Gesture Recognition in Robotic Surgery: a Review

Beatrice van Amsterdam, Matthew J. Clarkson, Danail Stoyanov

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments in IEEE Transactions on Biomedical Engineering, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05691 2021-01-29 cs.CV 57%

Learning Spatiotemporal Features via Video and Text Pair Discrimination

Tianhao Li, Limin Wang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11080 2021-01-28 cs.CV 57%

Deep Video Inpainting Detection

Peng Zhou, Ning Yu, Zuxuan Wu, Larry S. Davis, Abhinav Shrivastava, Ser-Nam Lim

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13318 2020-12-25 cs.CV 57%

Person Re-Identification using Deep Learning Networks: A Systematic Review

Ankit Yadav, Dinesh Kumar Vishwakarma

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 34 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12432 2020-12-25 eess.IV cs.CV 57%

Multi-Contrast Computed Tomography Healthy Kidney Atlas

Ho Hin Lee, Yucheng Tang, Kaiwen Xu, Shunxing Bao, Agnes B. Fogo, Raymond Harris, Mark P. de Caestecker, Mattias Heinrich, Jeffrey M. Spraggins, Yuankai Huo, Bennett A. Landman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07711 2020-12-21 cs.CV cs.LG 57%

Knowledge Distillation for Action Anticipation via Label Smoothing

Guglielmo Camporese, Pasquale Coscia, Antonino Furnari, Giovanni Maria Farinella, Lamberto Ballan

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.05134 2020-12-02 cs.RO cs.AI cs.LG 57%

Deep Imitation Learning for Bimanual Robotic Manipulation

Fan Xie, Alexander Chowdhury, M. Clara De Paolis Kaluza, Linfeng Zhao, Lawson L. S. Wong, Rose Yu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14618 2020-12-01 cs.IR cs.CL cs.SI 57%

CovidExplorer: A Multi-faceted AI-based Search and Visualization Engine for COVID-19 Information

Heer Ambavi, Kavita Vaishnaw, Udit Vyas, Abhisht Tiwari, Mayank Singh

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL

Comments 4 pages, 7 figures, The associated system can be accessed at http://covidexplorer.in, To be published in the Proceedings of the 29th ACM International Conference on Information and Knowledge Management (CIKM '20) (October 19-23, 2020)(Virtual Event, Ireland)

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06852 2020-11-16 cs.CV 57%

Discriminative Feature Representation with Spatio-temporal Cues for Vehicle Re-identification

J. Tu, C. Chen, X. Huang, J. He, X. Guan

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04258 2020-11-10 cs.CV 57%

Improved Soccer Action Spotting using both Audio and Video Streams

Bastien Vanderplaetse, Stéphane Dupont

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2020, pp. 896-897

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14705 2020-10-29 cs.CV 57%

Quantified Facial Temporal-Expressiveness Dynamics for Affect Analysis

Md Taufeeq Uddin, Shaun Canavan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 25th International Conference on Pattern Recognition (ICPR2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12662 2020-10-27 cs.MM cs.LG 57%

Short Video-based Advertisements Evaluation System: Self-Organizing Learning Approach

Yunjie Zhang, Fei Tao, Xudong Liu, Runze Su, Xiaorong Mei, Weicong Ding, Zhichen Zhao, Lei Yuan, Ji Liu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments Submitting to ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11838 2020-10-23 cs.CV 57%

Blind Video Temporal Consistency via Deep Video Prior

Chenyang Lei, Yazhou Xing, Qifeng Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2020; github link: github.com/ChenyangLEI/deep-video-prior

详情

展开后加载摘要…

URL PDF HTML 收藏