arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9150 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9150 篇

2302.01676 2023-02-07 cs.LG cs.AI cs.CL cs.CV cs.NE 82%

Show me your NFT and I tell you how it will perform: Multimodal representation learning for NFT selling price prediction

Davide Costa, Lucio La Cava, Andrea Tagarelli

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted paper at The ACM Web Conference 2023, April 30--May 04, 2023, Austin, Texas, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03238 2023-01-10 cs.CL cs.AI cs.LG cs.SD eess.AS 82%

MAQA: A Multimodal QA Benchmark for Negation

Judith Yue Li, Aren Jansen, Qingqing Huang, Joonseok Lee, Ravi Ganti, Dima Kuzmin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments NeurIPS 2022 SyntheticData4ML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06078 2022-12-07 cs.CY stat.ML 82%

MUTLA: A Large-Scale Dataset for Multimodal Teaching and Learning Analytics

Fangli Xu, Lingfei Wu, KP Thai, Carol Hsu, Wei Wang, Richard Tong

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

Comments 3 pages, 1 figure, 2 tables workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02580 2022-11-07 cs.CL cs.AI cs.CV 82%

Evaluating and Improving Factuality in Multimodal Abstractive Summarization

David Wan, Mohit Bansal

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2022 (17 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.11645 2022-10-19 stat.CO cs.LG eess.SP math.OC 82%

Cyclical Variational Bayes Monte Carlo for Efficient Multi-Modal Posterior Distributions Evaluation

Felipe Igea, Alice Cicirello

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments Accepted version in MSSP

Journal ref Mechanical Systems and Signal Processing, Volume 186, 2023, 109868, ISSN 0888-3270

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05297 2022-09-21 cs.CV cs.CL cs.GR cs.LG cs.MM 82%

BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis

Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, Bo Zheng

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 28 pages, 15 figures, Accepted by ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06416 2022-09-15 cs.CL cs.AI cs.MM 82%

ImageArg: A Multi-modal Tweet Dataset for Image Persuasiveness Mining

Zhexiong Liu, Meiqi Guo, Yue Dai, Diane Litman

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments In Argument Mining Workshop, held in conjunction with the International Conference on Computational Linguistics (COLING), October 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14087 2022-09-14 cs.MM cs.CL cs.CV 82%

CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation

Hao Sun, Hongyi Wang, Jiaqing Liu, Yen-Wei Chen, Lanfen Lin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11066 2022-08-24 cs.NE 82%

Enhanced Opposition Differential Evolution Algorithm for Multimodal Optimization

Shatendra Singh, Aruna Tiwari

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.13445 2022-05-27 cs.CV cs.AI cs.CL cs.IT cs.LG math.IT 82%

Mutual Information Divergence: A Unified Metric for Multimodal Generative Models

Jin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo, Sang-Woo Lee

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12617 2022-05-26 cs.CL cs.AI cs.CV 82%

DisinfoMeme: A Multimodal Dataset for Detecting Meme Intentionally Spreading Out Disinformation

Jingnong Qu, Liunian Harold Li, Jieyu Zhao, Sunipa Dev, Kai-Wei Chang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11588 2022-04-26 cs.IR cs.AI cs.CL cs.CV cs.LG 82%

Ad Creative Discontinuation Prediction with Multi-Modal Multi-Task Neural Survival Networks

Shunsuke Kitada, Hitoshi Iyatomi, Yoshifumi Seki

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 23 pages, 5 figures. Accepted by Appl. Sci. on March 29th, 2022

Journal ref Appl. Sci. 2022, 12(7), 3594

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09792 2022-03-01 eess.SY cs.RO cs.SY 82%

Stochastic MPC with Multi-modal Predictions for Traffic Intersections

Siddharth H. Nair, Vijay Govindarajan, Theresa Lin, Chris Meissen, H. Eric Tseng, Francesco Borrelli

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments Extended version of ITSC 2022 submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03242 2022-02-08 cs.LG stat.ML 82%

Unsupervised physics-informed disentanglement of multimodal data for high-throughput scientific discovery

Nathaniel Trask, Carianne Martinez, Kookjin Lee, Brad Boyce

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12307 2021-09-28 cs.CV cs.AI cs.MM 82%

Multi-Modal Multi-Instance Learning for Retinal Disease Recognition

Xirong Li, Yang Zhou, Jie Wang, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACM Multimedia 2021 (Main Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.11526 2021-09-24 cs.CV cs.CL cs.CY cs.LG cs.MM 82%

MARMOT: A Deep Learning Framework for Constructing Multimodal Representations for Vision-and-Language Tasks

Patrick Y. Wu, Walter R. Mebane

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 57 pages, 16 figures. Forthcoming in Computational Communication Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02993 2021-09-08 cs.CV cs.MM cs.SD eess.AS eess.IV 82%

Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors

Hasam Khalid, Minha Kim, Shahroz Tariq, Simon S. Woo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments 2 Figures, 2 Tables, Accepted for publication at the 1st Workshop on Synthetic Multimedia - Audiovisual Deepfake Generation and Detection (ADGD '21) at ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09889 2021-06-21 cs.CL cs.CV cs.MM 82%

GEM: A General Evaluation Benchmark for Multimodal Tasks

Lin Su, Nan Duan, Edward Cui, Lei Ji, Chenfei Wu, Huaishao Luo, Yongfei Liu, Ming Zhong, Taroon Bharti, Arun Sacheti

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by Findings of ACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00229 2021-02-05 cs.CL cs.AI cs.CV 82%

Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation

Wanrong Zhu, Xin Eric Wang, Tsu-Jui Fu, An Yan, Pradyumna Narayana, Kazoo Sone, Sugato Basu, William Yang Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.03678 2020-12-08 cs.CV cs.AI cs.CL cs.LG 82%

Generating Natural Questions from Images for Multimodal Assistants

Alkesh Patel, Akanksha Bindal, Hadas Kotek, Christopher Klein, Jason Williams

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 4 pages, 1 reference page, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11286 2020-11-24 cs.MM cs.AI cs.CV 82%

MEG: Multi-Evidence GNN for Multimodal Semantic Forensics

Ekraam Sabir, Ayush Jaiswal, Wael AbdAlmageed, Prem Natarajan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments To be published at ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06671 2020-10-15 cs.CL cs.AI cs.CV 82%

A Multi-Modal Method for Satire Detection using Textual and Visual Cues

Lily Li, Or Levi, Pedram Hosseini, David A. Broniatowski

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to the Third Workshop on NLP for Internet Freedom (NLP4IF): Censorship, Disinformation, and Propaganda. Co-located with COLING 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06442 2019-11-11 q-bio.QM cs.LG stat.ML 82%

Co-Attentive Cross-Modal Deep Learning for Medical Evidence Synthesis and Decision Making

Devin Taylor, Simeon Spasov, Pietro Liò

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract)

Comments 7 pages, 2 figures, Machine Learning for Health (ML4H) at NeurIPS 2019 - Extended Abstract, clarified graph and math notation, typos corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.04975 2019-04-18 q-bio.NC 82%

Multimodal Cross-registration and Quantification of Metric Distortions in Whole Brain Histology of Marmoset using Diffeomorphic Mappings

Brian C. Lee, Meng Kuan Lin, Yan Fu, Junichi Hata, Michael I. Miller, Partha P. Mitra

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.07478 2019-03-19 q-bio.QM 82%

Neurovascular coupling: insights from multi-modal dynamic causal modelling of fMRI and MEG

Amirhossein Jafarian, Vladimir Litvak, Hayriye Cagnan, Karl J. Friston, Peter Zeidman

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06686 2018-08-22 cs.MM cs.AI cs.CV cs.LG cs.SI 82%

Deep Multimodal Image-Repurposing Detection

Ekraam Sabir, Wael AbdAlmageed, Yue Wu, Prem Natarajan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments To be published at ACM Multimeda 2018 (orals)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.02356 2018-05-08 cs.CL cs.AI cs.IR cs.MA cs.MM 82%

Multimodal Machine Translation with Reinforcement Learning

Xin Qian, Ziyi Zhong, Jieli Zhou

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.08251 2017-04-21 cs.CL cs.AI cs.CV 82%

Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation

Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios P. Spithourakis, Lucy Vanderwende

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.04517 2017-04-18 cs.CL cs.AI cs.CV 82%

ShapeWorld - A new test methodology for multimodal language understanding

Alexander Kuhnle, Ann Copestake

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.01847 2016-09-27 cs.CV cs.AI cs.CL 82%

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, Marcus Rohrbach

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2016

详情

展开后加载摘要…

URL PDF HTML 收藏