arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2103.10139 2021-03-19 cs.CV 79%

Learning Multimodal Affinities for Textual Editing in Images

Or Perel, Oron Anschel, Omri Ben-Eliezer, Shai Mazor, Hadar Averbuch-Elor

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ACM Transactions on Graphics 2021, to be presented in SIGGRAPH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12338 2021-02-01 cs.RO cs.AI 79%

Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Ting Han, Sina Zarrieß

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments The 2nd Workshop on NLG for HRI colocated with The 13th International Conference on Natural Language Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04776 2020-12-10 cs.LG cs.AI 79%

A Data-Driven Analytical Framework of Estimating Multimodal Travel Demand Patterns using Mobile Device Location Data

Chenfeng Xiong, Aref Darzi, Yixuan Pan, Sepehr Ghader, Lei Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09067 2020-10-20 cs.CV 79%

Multimodal semantic forecasting based on conditional generation of future features

Kristijan Fugošić, Josip Šarić, Siniša Šegvić

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to German Conference on Pattern Recognition 2020. 24 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12429 2020-04-28 cs.CL 79%

Towards Multimodal Response Generation with Exemplar Augmentation and Curriculum Optimization

Zeyang Lei, Zekang Li, Jinchao Zhang, Fandong Meng, Yang Feng, Yujiu Yang, Cheng Niu, Jie Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00431 2020-03-03 cs.AI 79%

A Study on Multimodal and Interactive Explanations for Visual Question Answering

Kamran Alipour, Jurgen P. Schulze, Yi Yao, Avi Ziskind, Giedrius Burachas

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments http://ceur-ws.org/Vol-2560/paper44.pdf

Journal ref Proceedings of the Workshop on Artificial Intelligence Safety (SafeAI 2020) co-located with 34th AAAI Conference on Artificial Intelligence (AAAI 2020), New York, USA, Feb 7, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06484 2020-02-18 cs.CL 79%

A Multimodal Dialogue System for Conversational Image Editing

Tzu-Hsiang Lin, Trung Bui, Doo Soon Kim, Jean Oh

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at 2nd Conversational AI Workshop at NeurIPS 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.09792 2019-11-20 cs.CL 79%

A Multi-Modal Chinese Poetry Generation Model

Dayiheng Liu, Quan Guo, Wubo Li, Jiancheng Lv

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06663 2019-11-18 cs.LG cs.CV stat.ML 79%

MMGAN: Generative Adversarial Networks for Multi-Modal Distributions

Teodora Pandeva, Matthias Schubert

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04888 2019-06-13 cs.RO cs.CV cs.SY eess.SY 79%

Adaptive Navigation Scheme for Optimal Deep-Sea Localization Using Multimodal Perception Cues

Arturo Gomez Chavez, Qingwen Xu, Christian A. Mueller, Sören Schwertfeger, Andreas Birk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to IROS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13443 2019-06-03 cs.CL 79%

Symbol Emergence as an Interpersonal Multimodal Categorization

Yoshinobu Hagiwara, Hiroyoshi Kobayashi, Akira Taniguchi, Tadahiro Taniguchi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.13304 2019-04-25 cs.CV 79%

Acute and sub-acute stroke lesion segmentation from multimodal MRI

Albert Clèrigues, Sergi Valverde, Jose Bernal, Jordi Freixenet, Arnau Oliver, Xavier Lladó

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01735 2019-04-04 cs.CL 79%

Multi-Modal Generative Adversarial Network for Short Product Title Generation in Mobile E-Commerce

Jian-Guo Zhang, Pengcheng Zou, Zhao Li, Yao Wan, Xiuming Pan, Yu Gong, Philip S. Yu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by NAACL-HLT 2019. arXiv admin note: substantial text overlap with arXiv:1811.04498

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11955 2018-11-22 cs.CL 79%

Improving Context Modelling in Multimodal Dialogue Generation

Shubham Agarwal, Ondrej Dusek, Ioannis Konstas, Verena Rieser

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Journal ref Proceedings of the 11th International Conference on Natural Language Generation, pages 129-134, Tilburg, The Netherlands, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04498 2018-11-13 cs.CL 79%

Product Title Refinement via Multi-Modal Generative Adversarial Learning

Jianguo Zhang, Pengcheng Zou, Zhao Li, Yao Wan, Ye Liu, Xiuming Pan, Yu Gong, Philip S. Yu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Workshop on Visually Grounded Interaction and Language, NIPS, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03257 2018-05-10 cs.CL 79%

Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog

Jiaping Zhang, Tiancheng Zhao, Zhou Yu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00410 2018-04-03 cs.CV 79%

SyncGAN: Synchronize the Latent Space of Cross-modal Generative Adversarial Networks

Wen-Cheng Chen, Chien-Wen Chen, Min-Chun Hu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments 9 pages, Part of this work is accepted by IEEE International Conference on Multimedia Expo 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.11550 2018-04-02 cs.CV 79%

Multi-modal Disease Classification in Incomplete Datasets Using Geometric Matrix Completion

Gerome Vivar, Andreas Zwergal, Nassir Navab, Seyed-Ahmad Ahmadi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.06151 2017-09-20 cs.CV q-bio.NC q-bio.QM 79%

Multi-modal analysis of genetically-related subjects using SIFT descriptors in brain MRI

Kuldeep Kumar, Laurent Chauvin, Mathew Toews, Olivier Colliot, Christian Desrosiers

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Journal ref Proc. Computational Diffusion MRI, MICCAI Workshop, Québec City, Canada, September 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02337 2017-06-09 cs.CV cs.LG 79%

Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Network

Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, C. Lee Giles

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2017 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.02841 2017-04-11 cs.HC cs.CL 79%

From Modal to Multimodal Ambiguities: a Classification Approach

Maria Chiara Caschera, Fernando Ferri, Patrizia Grifoni

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 23 pages

Journal ref JNIT (Journal of Next Generation Information Technology), Volume 4 Issue 5, July, 2013,Pages 87-109, ISSN 2092-8637. GlobalCIS (Convergence Information Society, Republic of Korea)

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.02621 2017-04-11 cs.AI stat.ML 79%

Mixed Graphical Models for Causal Analysis of Multi-modal Variables

Andrew J Sedgewick, Joseph D. Ramsey, Peter Spirtes, Clark Glymour, Panayiotis V. Benos

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08472 2016-11-28 cs.CV 79%

Multimodal Latent Variable Analysis

Vardan Papyan, Ronen Talmon

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.01801 2016-08-29 cs.CV cs.LG stat.ML 79%

Variational methods for Conditional Multimodal Deep Learning

Gaurav Pandey, Ambedkar Dukkipati

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.07481 2016-06-27 cs.CL 79%

CUNI System for WMT16 Automatic Post-Editing and Multimodal Translation Tasks

Jindřich Libovický, Jindřich Helcl, Marek Tlustý, Pavel Pecina, Ondřej Bojar

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to the First Conference of Machine Translation (WMT16)

详情

展开后加载摘要…

URL PDF HTML 收藏
1003.1458 2010-03-09 cs.CR cs.CV 79%

Secured Cryptographic Key Generation From Multimodal Biometrics: Feature Level Fusion of Fingerprint and Iris

A. Jagadeesan, K. Duraiswamy

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Pages IEEE format, International Journal of Computer Science and Information Security, IJCSIS February 2010, ISSN 1947 5500, http://sites.google.com/site/ijcsis/

Journal ref International Journal of Computer Science and Information Security, IJCSIS, Vol. 7, No. 2, pp. 028-037, February 2010, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11136 2025-12-10 hep-ph hep-ex hep-th nucl-ex nucl-th 79%

Heavy-flavor multimodal fragmentation to $S$-wave pentacharms at next-generation hadron colliders

重味多模碎片化到下一代强子对撞机的S波五重态

Francesco Giovanni Celiberto

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 研究了在下一代强子对撞机中通过多模碎片化产生S波五重态的机制及现象学影响。

Comments 49 pages, 9 figures, 245 references, published in Eur. Phys. J. C. One novel set of multimodal (direct multicharm and diquark-like initial-scale inputs) "PentaQuarks with 5 heavy Quarks" (PQ5Q1.0) NLO collinear fragmentation functions released in LHAPDF format and publicly available from https://github.com/FGCeliberto/Collinear_FFs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21835 2025-10-28 cs.LG cs.AI cs.CL cs.CV 79%

A Multimodal, Multitask System for Generating E Commerce Text Listings from Images

Nayan Kumar Singh

专题命中 多模态生成 :multimodal(title,comments);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 10 figures, 11 tables. Code can be found at: https://github.com/SinghNayanKumar/multimodal-product-lister/

详情

展开后加载摘要…

URL PDF HTML 收藏
0708.3575 2009-12-01 cs.HC 79%

How really effective are Multimodal Hints in enhancing Visual Target Spotting? Some evidence from a usability study

Suzanne Kieffer, Noëlle Carbonell

专题命中 多模态生成 :multimodal(title,abstract)

Comments 9 pages

Journal ref Journal on Multimodal Interaction (JMUI), 1 (2007) 1-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18504 2026-07-27 cs.LG cs.AI cs.CV 版本更新 79%

Now We Know? A Systematic Comparison of TerraMind and THOR

我们现在知道了吗?TerraMind和THOR的系统比较

Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-Børre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci

机构 * University of Münster(明斯特大学) IBM Research(IBM研究院) Norwegian Computing Center(挪威计算中心)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.CV、cs.AI

AI总结 通过对TerraMind和THOR两个地理空间基础模型对比,研究其在补丁大小、解码器复杂性等方面的差异轴,发现架构设计选择对性能差异影响更大,体现互补投资策略,还得出假设和诊断消融方法,有望推广到未来模型。

详情

展开后加载摘要…

URL PDF HTML 收藏