MultiMAE: Multi-modal Multi-task Masked Autoencoders
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Project page at https://multimae.epfl.ch
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Project page at https://multimae.epfl.ch
专题命中 多模态评测 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL
Comments ACL 2022 Main conference (Long Paper)
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
Comments Accepted to Findings of ACL 2021
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 6 pages, 4 figures, 1 table, accepted to BVM 2022
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments Accepted for publication in AAAI-2022
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments In Submission
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments Accepted in Neurocomputing
专题命中 多模态评测 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV
Comments 8 pages
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Accepted at Perception Beyond Visible Spectrum Workshop, CVPR 2019
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Journal ref IEEE Transactions on Medical Imaging, vol. 39, no. 5, pp. 1703-1711, May 2020
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted at CVPR 2020. For a demo video, see http://tiny.cc/xmuda
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL
Comments 6 pages, 3 figures
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted to CVPR 2019. Project Page: http://visgel.csail.mit.edu/
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Working Draft
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
M$^3$Eval: 通过认知基础视频任务的多模态记忆评估
机构 * School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) ; Yuanpei College, Peking University(北京大学元培学院) ; Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) ; School of Psychological and Cognitive Sciences, Peking University(北京大学心理学与认知科学学院) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 提出首个多模态模型记忆评估框架M$^3$Eval,通过认知心理学设计的视频任务系统评估模型在记忆保持、忠实性和鲁棒性上的表现,发现模型在并行视频流处理、干扰模式、时空记忆和符号记忆方面的显著缺陷。
Comments We present an evaluation designed for multi-modal memory in multi-modal models
测试时匹配:在多模态模型中解锁组合推理
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出测试时匹配算法,通过改进评估指标提升多模态模型的组合推理能力,使SigLIP-B16和GPT-4.1在Winoground等基准上取得新突破。
Comments To appear at ICLR 2026; extended results to generative multimodal models
机构 * Fontys University of Applied Sciences(Fontys应用科学大学) ; Technical University of Eindhoven(埃因霍温技术大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models for biodiversity identification using images, language and spatiotemporal data
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments LogicVista benchmarks the logical reasoning of multimodal large language models in visual tasks
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Multimodal Benchmark, Project Url: https://zeyofu.github.io/blink/, ECCV 2024
PanDent:面向牙科放射学中全面的牙级结构-语言一致性
机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院) ; Imperial College London(帝国理工学院) ; University of Science and Technology of China(中国科学技术大学) ; Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) ; Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM
AI总结 本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。
十年视觉语言人工智能模型中准确性和视觉认知错误的演变
机构 * Psychological & Brain Sciences, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校心理与脑科学系) ; Department of Computer Science, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校计算机科学系) ; Department of Electrical and Computer Engineering, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校电气与计算机工程系)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 研究十年间视觉语言模型进展,引入CSB数据集,评估模型在其上及MS-COCO样本中的准确性与视觉认知错误类型,发现MLLM消除简单与复杂场景描述准确性差距,几乎消除多数错误类型,为模型发展提供全面评估。
解释比单独预测更难:评估基于概念的MLLM解释作为ICL视觉分类器
专题命中 多模态评测 :MLLM(title_cn,abstract_cn);multimodal(abstract);分类 cs.CL、cs.AI
AI总结 本文通过五种形式化程度递增的条件,系统评估多模态大语言模型在少样本上下文学习中的基于概念的可解释性,发现解释比预测更难,且强制生成形式化解释会降低预测准确性。
Comments Accepted to the CompLearn Workshop at ICML 2026
VLRS-Bench: 一种面向遥感的视觉-语言推理基准
机构 * School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出VLRS-Bench,首个专注于复杂遥感推理的基准,包含2000个问题-答案对,涵盖14项任务和八个时间阶段,揭示现有MLLM在遥感任务中的瓶颈。
PPU-Bench: 用于视觉语言模型个性化部分遗忘的现实世界基准
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Pengcheng Laboratory(鹏城实验室) ; The Hong Kong Polytechnic University(香港理工大学) ; Sichuan University(四川大学) ; Zhejiang Normal University(浙江师范大学)
专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出PPU-Bench,一个无需微调的现实世界基准,用于评估视觉语言模型中个性化部分遗忘的效果,通过24K多模态和单模态样本测试遗忘与保留的平衡及跨模态一致性。
专题命中 多模态评测 :multimodal(abstract,comments);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Total 120 pages. See our project at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models
GeoExplain:基于街景图像视觉信息层次的多模态推理
机构 * The University of Queensland(昆士兰大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
AI总结 针对现有多模态推理任务缺乏对不同粒度视觉线索推理的研究缺口,提出GeoExplain数据集与SightSense方法,实现街景图像的地理定位预测与可解释性说明,且方法在该数据集上表现优异。
Comments Updated version