arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9170 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9170 篇

2205.10745 2022-05-24 cs.CV cs.AI 81%

Classification of Quasars, Galaxies, and Stars in the Mapping of the Universe Multi-modal Deep Learning

Sabeesh Ethiraj, Bharath Kumar Bolla

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Presented at Deep Learning Developers Conference, 2021, Bangalore

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08007 2022-05-18 cs.MM cs.SD eess.AS eess.IV 81%

Perceptual Evaluation on Audio-visual Dataset of 360 Content

Randy F Fela, Andréas Pastor, Patrick Le Callet, Nick Zacharov, Toinon Vigier, Søren Forchhammer

专题命中 多模态评测 :audio-visual(title);multimodal(abstract);分类 cs.MM、eess.AS

Comments 6 pages, 5 figures, International Conference on Multimedia and Expo 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01133 2022-05-09 cs.CL cs.CV cs.LG 81%

Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation

Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Hassan Muhammad, Ibrahim Sa'id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at Language Resources and Evaluation Conference 2022 (LREC2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01850 2022-05-05 cs.CL cs.CV 81%

Visual Commonsense in Pretrained Unimodal and Multimodal Models

Chenyu Zhang, Benjamin Van Durme, Zhuowan Li, Elias Stengel-Eskin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments To appear in NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01879 2022-05-05 cs.CV cs.AI 81%

Learning Two-Stream CNN for Multi-Modal Age-related Macular Degeneration Categorization

Weisen Wang, Xirong Li, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (J-BHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07154 2022-05-03 cs.CL cs.CV 81%

MMChat: Multi-Modal Chat Dataset on Social Media

Yinhe Zheng, Guanyi Chen, Xin Liu, Jian Sun

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by LREC2022. Dataset available in https://github.com/silverriver/MMChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09788 2022-04-22 cs.CV cs.AI 81%

SELMA: SEmantic Large-scale Multimodal Acquisitions in Variable Weather, Daytime and Viewpoints

Paolo Testolina, Francesco Barbato, Umberto Michieli, Marco Giordani, Pietro Zanuttigh, Michele Zorzi

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 figures, 14 tables. This paper has been submitted to IEEE. Copyright may change without notice

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15125 2022-04-06 cs.CV cs.CL cs.LG 81%

Text2Pos: Text-to-Point-Cloud Cross-Modal Localization

Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixe

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR2022 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00061 2022-03-22 cs.CV cs.CL cs.CY cs.LG 81%

Open-Domain, Content-based, Multi-modal Fact-checking of Out-of-Context Images via Online Resources

Sahar Abdelnabi, Rakibul Hasan, Mario Fritz

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09824 2022-03-21 cs.CV cs.LG eess.AS 81%

Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?

Cho-Ying Wu, Chin-Cheng Hsu, Ulrich Neumann

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted to CVPR 2022. Project page: https://choyingw.github.io/works/Voice2Mesh/index.html. This version supersedes arXiv:2104.10299

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06419 2022-03-15 cs.CL cs.AI 81%

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

Shivani Kumar, Atharva Kulkarni, Md Shad Akhtar, Tanmoy Chakraborty

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted in ACL 2022. 13 pages, 4 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04990 2022-01-14 cs.LG cs.AI cs.CV 81%

Toddler-Guidance Learning: Impacts of Critical Period on Multimodal AI Agents

Junseok Park, Kwanyoung Park, Hyunseok Oh, Ganghun Lee, Minsu Lee, Youngki Lee, Byoung-Tak Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICMI2021 Oral Presentation, 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03879 2021-12-23 cs.CL cs.CV 81%

AI2D-RST: A multimodal corpus of 1000 primary school science diagrams

Tuomo Hiippala, Malihe Alikhani, Jonas Haverinen, Timo Kalliokoski, Evanfiya Logacheva, Serafina Orekhova, Aino Tuomainen, Matthew Stone, John A. Bateman

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 24 pages; revised version submitted to Language Resources & Evaluation

Journal ref Language Resources and Evaluation 55(3), 2021, pp. 661-688

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03521 2021-12-08 cs.CL cs.AI 81%

UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

Yichen Huang, Yuchen Wang, Yik-Cheung Tam

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00800 2021-12-03 cs.CL cs.AI 81%

Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text

Christopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Jonghyun Choi, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments In EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07979 2021-11-16 cs.SD cs.AI cs.LG cs.SY eess.AS eess.SY q-bio.NC 81%

Metric-based multimodal meta-learning for human movement identification via footstep recognition

Muhammad Shakeel, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.08092 2021-10-29 cs.LG cs.AI cs.CV 81%

An AutoML-based Approach to Multimodal Image Sentiment Analysis

Vasco Lopes, António Gaspar, Luís A. Alexandre, João Cordeiro

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11899 2021-10-25 cs.CV cs.CL 81%

Challenges in Procedural Multimodal Machine Comprehension:A Novel Way To Benchmark

Pritish Sahu, Karan Sikka, Ajay Divakaran

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07235 2021-10-22 cs.CV cs.AI cs.LG 81%

HUMAN4D: A Human-Centric Multimodal Dataset for Motions and Immersive Media

Anargyros Chatzitofis, Leonidas Saroglou, Prodromos Boutis, Petros Drakoulis, Nikolaos Zioulis, Shishir Subramanyam, Bart Kevelham, Caecilia Charbonnier, Pablo Cesar, Dimitrios Zarpalas, Stefanos Kollias, Petros Daras

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref IEEE Access, 8, 176241-176262, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08667 2021-10-22 cs.CL cs.AI 81%

SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations

Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12212 2021-10-01 cs.CL cs.CV cs.CY 81%

An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog

Xingyao Wang, David Jurgens

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Findings of EMNLP 2021; 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13086 2021-09-28 cs.CV cs.AI 81%

MFEViT: A Robust Lightweight Transformer-based Network for Multimodal 2D+3D Facial Expression Recognition

Hanting Li, Mingzhe Sui, Zhaoqing Zhu, Feng Zhao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9pages,6 figures,5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01656 2021-09-28 cs.CL cs.AI 81%

IITP at WAT 2021: System description for English-Hindi Multimodal Translation Task

Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09487 2021-09-21 cs.CV cs.AI cs.LG 81%

Dyadformer: A Multi-modal Transformer for Long-Range Modeling of Dyadic Interactions

David Curto, Albert Clapés, Javier Selva, Sorina Smeureanu, Julio C. S. Jacques Junior, David Gallardo-Pujol, Georgina Guilera, David Leiva, Thomas B. Moeslund, Sergio Escalera, Cristina Palmero

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to the 2021 ICCV Workshop on Understanding Social Behavior in Dyadic and Small Group Interactions

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01839 2021-09-07 cs.CL cs.CV 81%

Towards Expressive Communication with Internet Memes: A New Multimodal Conversation Dataset and Benchmark

Zhengcong Fei, Zekang Li, Jinchao Zhang, Yang Feng, Jie Zhou

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03149 2021-09-02 cs.CV cs.AI 81%

Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering

Corentin Dancette, Remi Cadene, Damien Teney, Matthieu Cord

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2021. Code is available at https://github.com/cdancette/detect-shortcuts

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08920 2021-08-24 cs.LG cs.AI cs.CV 81%

Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach

Chuanbo Hu, Minglei Yin, Bin Liu, Xin Li, Yanfang Ye

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by CIKM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01453 2021-08-04 cs.IR cs.AI cs.CL 81%

PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling

Xiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song, Hao Zhang, Jindong Chen

专题命中 多模态评测 :image-text(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.07956 2021-07-19 cs.SD cs.CL eess.AS 81%

A Multimodal Machine Learning Framework for Teacher Vocal Delivery Evaluation

Hang Li, Yu Kang, Yang Hao, Wenbiao Ding, Zhongqin Wu, Zitao Liu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments AIED'21: The 22nd International Conference on Artificial Intelligence in Education, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13213 2021-06-25 cs.LG cs.AI cs.CL cs.HC 81%

Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data

Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2021. arXiv admin note: substantial text overlap with arXiv:2012.02359

详情

展开后加载摘要…

URL PDF HTML 收藏