arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9213 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9213 篇

2107.12719 2022-04-22 cs.MM cs.CV cs.SD eess.AS 67%

The CORSMAL benchmark for the prediction of the properties of containers

Alessio Xompero, Santiago Donaher, Vladimir Iashin, Francesca Palermo, Gökhan Solak, Claudio Coppola, Reina Ishikawa, Yuichi Nagao, Ryo Hachiuma, Qi Liu, Fan Feng, Chuanlin Lan, Rosa H. M. Chan, Guilherme Christmann, Jyun-Ting Song, Gonuguntla Neeharika, Chinnakotla Krishna Teja Reddy, Dinesh Jain, Bakhtawar Ur Rehman, Andrea Cavallaro

专题命中 多模态评测 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Authors' post-print accepted for publication in IEEE Access, see https://doi.org/10.1109/ACCESS.2022.3166906 . 14 pages, 6 tables, 7 figures

Journal ref IEEE Access, vol. 10, 2022, 1-15

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.14795 2022-03-17 cs.LG cs.CL cs.CV cs.SD eess.AS 67%

Perceiver IO: A General Architecture for Structured Inputs & Outputs

Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, Joāo Carreira

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments ICLR 2022 camera ready. Code: https://dpmd.ai/perceiver-code

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11589 2022-03-17 cs.CV cs.AI cs.CL cs.LG cs.RO 67%

VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

Ayush Shrivastava, Karthik Gopalakrishnan, Yang Liu, Robinson Piramuthu, Gokhan Tür, Devi Parikh, Dilek Hakkani-Tür

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at Findings of the Annual Meeting of the Association for Computational Linguistics (ACL) 2022, previous version accepted at Visually Grounded Interaction and Language (ViGIL) Workshop at NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02986 2022-03-08 cs.CV cs.AI cs.CL 67%

Modeling Coreference Relations in Visual Dialog

Mingxiao Li, Marie-Francine Moens

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.12141 2022-01-20 stat.ML cs.LG 67%

Stein Variational Gaussian Processes

Thomas Pinder, Christopher Nemeth, David Leslie

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract)

Comments 26 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10906 2021-11-19 cs.CV cs.AI cs.CL cs.LG 67%

Single-Modal Entropy based Active Learning for Visual Question Answering

Dong-Jin Kim, Jae Won Cho, Jinsoo Choi, Yunjae Jung, In So Kweon

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06860 2021-09-15 cs.CL cs.AI cs.CV 67%

Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning

Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, Kai-Wei Chang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2021. Code and data are available at https://github.com/WadeYin9712/GD-VCR

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.03744 2021-08-20 cs.CL cs.AI cs.CV 67%

e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations

Virginie Do, Oana-Maria Camburu, Zeynep Akata, Thomas Lukasiewicz

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref IEEE CVPR Workshop on Fair, Data Efficient and Trusted Computer Vision, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.00316 2021-08-03 cs.CV cs.AI cs.CL cs.LG 67%

Chest ImaGenome Dataset for Clinical Reasoning

Joy T. Wu, Nkechinyere N. Agu, Ismini Lourentzou, Arjun Sharma, Joseph A. Paguio, Jasper S. Yao, Edward C. Dee, William Mitchell, Satyananda Kashyap, Andrea Giovannini, Leo A. Celi, Mehdi Moradi

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Dataset available on PhysioNet (https://doi.org/10.13026/wv01-y230)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13054 2021-07-29 cs.AI cs.CL cs.CV cs.LG 67%

Exceeding the Limits of Visual-Linguistic Multi-Task Learning

Cameron R. Wolfe, Keld T. Lundgaard

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04884 2021-06-22 cs.CV cs.LG cs.MM cs.SD eess.AS 67%

Piano Skills Assessment

Paritosh Parmar, Jaiden Reddy, Brendan Morris

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Dataset is available from: https://github.com/ParitoshParmar/Piano-Skills-Assessment

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.02626 2021-05-07 cs.CV cs.AI cs.CL cs.LG 67%

A First Look: Towards Explainable TextVQA Models via Visual and Textual Explanations

Varun Nagaraj Rao, Xingjian Zhen, Karen Hovsepian, Mingwei Shen

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is done when Xingjian was an intern in Amazon PARS group, summer 2020. This paper is accepted by NAACL-MAI-Workshop, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.06732 2021-02-16 cs.CV cs.AI cs.IR cs.LG cs.MM 67%

Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution

Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, Mingxiang Cai

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 8 pages, 5 figures, to be published in AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.10210 2020-12-21 cs.CV cs.AI cs.CL 67%

On Modality Bias in the TVQA Dataset

Thomas Winterbottom, Sarah Xiao, Alistair McLean, Noura Al Moubayed

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 10 pages, 4 Figures, 2 Tables, +Supp Mats, BMVC 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00330 2020-11-19 cs.CV cs.AI cs.CL 67%

Visuo-Linguistic Question Answering (VLQA) Challenge

Shailaja Keyur Sampat, Yezhou Yang, Chitta Baral

专题命中 多模态评测 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Findings of EMNLP 2020 (22 pages, 13 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.13862 2020-09-30 cs.CV cs.AI cs.MM 67%

Where is the Model Looking At?--Concentrate and Explain the Network Attention

Wenjia Xu, Jiuniu Wang, Yang Wang, Guangluan Xu, Wei Dai, Yirong Wu

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 3, pp. 506-516, March 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.11401 2020-01-31 cs.HC eess.SP 67%

Visuohaptic augmented feedback for enhancing motor skills acquisition

Ali Asadipour, Kurt Debattista, Alan Chalmers

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract)

Comments 11 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08034 2020-01-23 cs.CL cs.AI cs.CV 67%

ManyModalQA: Modality Disambiguation and QA over Diverse Inputs

Darryl Hannan, Akshay Jain, Mohit Bansal

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2020 (10 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.06354 2020-01-20 cs.CL cs.AI cs.CV 67%

Modality-Balanced Models for Visual Dialogue

Hyounghun Kim, Hao Tan, Mohit Bansal

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2020 (11 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.02103 2019-11-07 cs.CV cs.CL cs.MM 67%

Recurrent Instance Segmentation using Sequences of Referring Expressions

Alba Herrera-Palacio, Carles Ventura, Carina Silberer, Ionut-Teodor Sorodoc, Gemma Boleda, Xavier Giro-i-Nieto

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments 3rd NeurIPS Workshop on Visually Grounded Interaction and Language (ViGIL, 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.03166 2019-09-20 cs.CV cs.AI cs.CL cs.LG 67%

CLEVR-Dialog: A Diagnostic Dataset for Multi-Round Reasoning in Visual Dialog

Satwik Kottur, José M. F. Moura, Devi Parikh, Dhruv Batra, Marcus Rohrbach

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 13 pages, 11 figures, 3 tables, accepted as a short paper at NAACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.09115 2019-04-22 cs.CV cs.HC cs.MM cs.SD eess.AS 67%

Listen to the Image

Di Hu, Dong Wang, Xuelong Li, Feiping Nie, Qi Wang

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by CVPR2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.01670 2018-07-06 cs.CL cs.AI cs.CV cs.LG 67%

Encoding Spatial Relations from Natural Language

Tiago Ramalho, Tomáš Kočiský, Frederic Besse, S. M. Ali Eslami, Gábor Melis, Fabio Viola, Phil Blunsom, Karl Moritz Hermann

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.01322 2018-05-15 cs.CL cs.AI cs.CV cs.LG 67%

Deep learning evaluation using deep linguistic processing

Alexander Kuhnle, Ann Copestake

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03017 2017-12-20 cs.CV cs.AI cs.CL stat.ML 67%

Learning Visual Reasoning Without Strong Priors

Ethan Perez, Harm de Vries, Florian Strub, Vincent Dumoulin, Aaron Courville

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Full AAAI 2018 paper is at arXiv:1709.07871. Presented at ICML 2017's Machine Learning in Speech and Language Processing Workshop. Code is at http://github.com/ethanjperez/film

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.11601 2017-11-03 cs.CL cs.AI cs.CV 67%

Whodunnit? Crime Drama as a Case for Natural Language Understanding

Lea Frermann, Shay B. Cohen, Mirella Lapata

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear in Transactions of the Association for Computational Linguistics (TACL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.09684 2017-09-15 cs.CV cs.AI cs.CL 67%

An Analysis of Visual Question Answering Algorithms

Kushal Kafle, Christopher Kanan

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear in ICCV 2017. Visit http://kushalkafle.com/projects/tdiuc to download the dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.02697 2016-11-28 cs.CV cs.AI cs.CL 67%

Ask Your Neurons: A Deep Learning Approach to Visual Question Answering

Mateusz Malinowski, Marcus Rohrbach, Mario Fritz

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Improved version, it also has a final table from the VQA challenge, and more baselines on DAQUAR

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.01121 2015-10-02 cs.CV cs.AI cs.CL 67%

Ask Your Neurons: A Neural-based Approach to Answering Questions about Images

Mateusz Malinowski, Marcus Rohrbach, Mario Fritz

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICCV'15 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08991 2026-08-12 cs.CV cs.AI 版本更新 66%

PinpointQA: A Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

PinpointQA: 一个用于室内视频中微小物体中心空间理解的数据集和基准

Zhiyu Zhou, Peilin Liu, Ruoxuan Zhang, Luyang Zhang, Cheng Zhang, Hongxia Xie, Wen-Huang Cheng

机构 * Jilin University(吉林大学) National Taiwan University(国立台湾大学)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV、cs.AI

AI总结 本文提出PinpointQA数据集,用于评估多模态大语言模型在室内视频中对微小物体空间定位的精准理解能力,通过四个递进任务验证模型性能,并展示其作为诊断基准和训练数据集的效果。

Comments Accepted at the ECCV 2026 Workshop on Embodied Multimodal Reasoning in Physical Environments (EMR)

详情

展开后加载摘要…

URL PDF HTML 收藏