BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
机构 * Baidu Inc.(百度公司)
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 10 pages
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Baidu Inc.(百度公司)
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 10 pages
专题命中 多模态生成 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.AI、cs.MM
专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL
Comments Accepted to ICLR 2025. Data and code are available at https://github.com/ChartMimic/ChartMimic
专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2403.16188
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.MM
Comments Accepted by Information Fusion. The code is available at https://github.com/zhangbo-nlp/VIKDF
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI
Comments Erratum: We identified an error in the calculation of the F1 score in table 4 reported in a previous version of this work. The performance of the new result is better than the previous one. The corrected values are included in this updated version of the paper. These changes do not alter the primary conclusions of our research
Journal ref MM '2024: Proceedings of the 32nd ACM International Conference on Multimedia, Pages 4814 - 4822
专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments Accepted in ACMMM 2024
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI
Comments Accepted by LREC-COLING 2024 as a short paper. ACL Anthology URL: [https://aclanthology.org/2024.lrec-main.1142/]
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments Accepted as Proceedings Paper at ML4H 2024
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted by NeurIPS-24
专题命中 多模态生成 :any-to-any(title,abstract);multi-modal(abstract);分类 cs.AI、cs.MM
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments Project page: https://sais-fuxi.github.io/projects/evalalign/
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL
Comments EMNLP 2024
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments ICML 2024. Project: https://github.com/YangLing0818/RPG-DiffusionMaster
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments Code: https://aka.ms/Kosmos-G Project Page: https://xichenpan.github.io/kosmosg
专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI
专题命中 多模态生成 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments NeurIPS 2023 Datasets and Benchmarks Track
专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI