arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9170 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9170 篇

2307.06775 2023-11-07 cs.LG cs.AI cs.CL cs.SI 81%

A Novel Site-Agnostic Multimodal Deep Learning Model to Identify Pro-Eating Disorder Content on Social Media

Jonathan Feldman

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13786 2023-11-01 cs.CV cs.AI cs.LG 81%

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Viorica Pătrăucean, Lucas Smaira, Ankush Gupta, Adrià Recasens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alex Frechette, Hanna Klimczak, Raphael Koster, Junlin Zhang, Stephanie Winkler, Yusuf Aytar, Simon Osindero, Dima Damen, Andrew Zisserman, João Carreira

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 37th Conference on Neural Information Processing Systems (NeurIPS 2023) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18046 2023-10-30 cs.CL cs.CV 81%

ViCLEVR: A Visual Reasoning Dataset and Hybrid Multimodal Fusion Model for Visual Question Answering in Vietnamese

Khiem Vinh Tran, Hao Phu Phan, Kiet Van Nguyen, Ngan Luu Thuy Nguyen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments A pre-print version and submitted to journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14605 2023-10-24 cs.CL cs.MM 81%

M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis

Fei Zhao, Chunhui Li, Zhen Wu, Yawen Ouyang, Jianbing Zhang, Xinyu Dai

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted by EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11465 2023-10-19 cs.LG cs.AI cs.CL 81%

BaitBuster-Bangla: A Comprehensive Dataset for Clickbait Detection in Bangla with Multi-Feature and Multi-Modal Analysis

Abdullah Al Imran, Md Sakib Hossain Shovon, M. F. Mridha

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10967 2023-10-18 cs.CL cs.AI cs.HC 81%

EXMODD: An EXplanatory Multimodal Open-Domain Dialogue dataset

Hang Yin, Pinren Lu, Ziang Li, Bin Sun, Kan Li

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06487 2023-10-17 cs.CV cs.AI 81%

Evaluating Explainable AI on a Multi-Modal Medical Imaging Task: Can Existing Algorithms Fulfill Clinical Requirements?

Weina Jin, Xiaoxiao Li, Ghassan Hamarneh

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09036 2023-10-16 cs.CL cs.MM 81%

MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks

Xiaocui Yang, Wenfang Wu, Shi Feng, Ming Wang, Daling Wang, Yang Li, Qi Sun, Yifei Zhang, Xiaoming Fu, Soujanya Poria

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Underview

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16772 2023-10-10 cs.CV cs.AI cs.RO 81%

XVO: Generalized Visual Odometry via Cross-Modal Self-Training

Lei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-Bar

专题命中 多模态评测 :cross-modal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2023, Paris https://genxvo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00845 2023-10-03 cs.CL cs.AI 81%

Application of frozen large-scale models to multimodal task-oriented dialogue

Tatsuki Kawamoto, Takuma Suzuki, Ko Miyama, Takumi Meguro, Tomohiro Takagi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03897 2023-10-03 cs.CL cs.CV 81%

Factify 2: A Multimodal Fake News and Satire News Dataset

S Suryavardan, Shreyash Mishra, Parth Patwa, Megha Chakraborty, Anku Rani, Aishwarya Reganti, Aman Chadha, Amitava Das, Amit Sheth, Manoj Chinnakotla, Asif Ekbal, Srijan Kumar

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Defactify2 @AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06495 2023-09-14 cs.CL cs.AI cs.PF 81%

AGIBench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark for Large Language Models

Fei Tang, Wanling Gao, Luzhou Peng, Jianfeng Zhan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01383 2023-09-06 cs.CV cs.AI 81%

LoRA-like Calibration for Multimodal Deception Detection using ATSFace Data

Shun-Wen Hsiao, Cheng-Yuan Sun

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12156 2023-08-24 cs.CV cs.AI 81%

Multimodal Latent Emotion Recognition from Micro-expression and Physiological Signals

Liangfei Zhang, Yifei Qian, Ognjen Arandjelovic, Anthony Zhu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07686 2023-08-16 cs.CV cs.AI 81%

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

Hong Li, Xingyu Li, Pengbo Hu, Yinuo Lei, Chunxiao Li, Yi Zhou

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02508 2023-08-08 eess.IV cs.AI cs.CV 81%

A Multimodal Supervised Machine Learning Approach for Satellite-based Wildfire Identification in Europe

Angelica Urbanelli, Luca Barco, Edoardo Arnaudo, Claudio Rossi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at IGARSS 2023, short paper (4 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00628 2023-08-08 cs.CV cs.AI cs.LG 81%

Human-M3: A Multi-view Multi-modal Dataset for 3D Human Pose Estimation in Outdoor Scenes

Bohao Fan, Siqi Wang, Wenxuan Guo, Wenzhao Zheng, Jianjiang Feng, Jie Zhou

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Code and data will be released on https://github.com/soullessrobot/Human-M3-Dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16125 2023-08-03 cs.CL cs.CV 81%

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, Ying Shan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Technical Report; Project released at: https://github.com/AILab-CVC/SEED-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11952 2023-07-25 cs.CV cs.AI 81%

Pathology-and-genomics Multimodal Transformer for Survival Outcome Prediction

Kexin Ding, Mu Zhou, Dimitris N. Metaxas, Shaoting Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to MICCAI2023 (Top14%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02716 2023-07-07 cs.CL cs.CV 81%

CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization

Min Xiao, Junnan Zhu, Haitao Lin, Yu Zhou, Chengqing Zong

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments acl2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02499 2023-07-07 cs.CL cs.AI 81%

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, Fei Huang

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18326 2023-07-04 cs.CV cs.AI 81%

BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine Translation

Liyan Kang, Luyang Huang, Ningxin Peng, Peihao Zhu, Zewei Sun, Shanbo Cheng, Mingxuan Wang, Degen Huang, Jinsong Su

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15977 2023-06-29 cs.CV cs.AI 81%

A Dimensional Structure based Knowledge Distillation Method for Cross-Modal Learning

Lingyu Si, Hongwei Dong, Wenwen Qiang, Junzhi Yu, Wenlong Zhai, Changwen Zheng, Fanjiang Xu, Fuchun Sun

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13968 2023-06-27 cs.CL cs.AI 81%

Fusing Multimodal Signals on Hyper-complex Space for Extreme Abstractive Text Summarization (TL;DR) of Scientific Contents

Yash Kumar Atri, Vikram Goyal, Tanmoy Chakraborty

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ADS-SIGKDD2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05493 2023-06-12 cs.CV cs.AI cs.LG 81%

Multi-Modal Classifiers for Open-Vocabulary Object Detection

Prannay Kaul, Weidi Xie, Andrew Zisserman

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICML 2023, project page: https://www.robots.ox.ac.uk/vgg/research/mm-ovod/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04738 2023-06-09 cs.CV cs.AI 81%

MultiEarth 2023 -- Multimodal Learning for Earth and Environment Workshop and Challenge

Miriam Cha, Gregory Angelides, Mark Hamilton, Andy Soszynski, Brandon Swenson, Nathaniel Maidel, Phillip Isola, Taylor Perron, Bill Freeman

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04021 2023-06-08 cs.CV cs.AI cs.LG cs.RO 81%

Energy-Based Models for Cross-Modal Localization using Convolutional Transformers

Alan Wu, Michael S. Ryoo

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18641 2023-05-31 cs.CL cs.CV 81%

Enhanced Chart Understanding in Vision and Language Task via Cross-modal Pre-training on Plot Table Pairs

Mingyang Zhou, Yi R. Fung, Long Chen, Christopher Thomas, Heng Ji, Shih-Fu Chang

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07044 2023-05-30 cs.CV cs.AI 81%

SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation

Yi Wang, Nassim Ait Ali Braham, Zhitong Xiong, Chenying Liu, Conrad M Albrecht, Xiao Xiang Zhu

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine. 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17388 2023-05-30 cs.CL cs.CV 81%

MPCHAT: Towards Multimodal Persona-Grounded Conversation

Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏