Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments 11 pages, 1 figure
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments 11 pages, 1 figure
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ; University of Pisa(比萨大学) ; IIT-CNR(意大利国家研究 council(IIT))
专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)
Comments ICCV 2025 Workshop on What is Next in Multimodal Foundation Models
机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)
专题命中 图文多模态 :multimodal(title);分类 cs.CV
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Engineering Research Center of Machine Learning and Industry Intelligence(机器学习与工业智能工程研究中心)
专题命中 图文多模态 :image-text(title);分类 cs.CV
机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗赖堡-亚历山大大学埃尔兰根-纽伦堡) ; Imperial College London(伦敦帝国理工学院)
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments 20 pages, 6 figures. To appear in Proc. MIDL 2025 (PMLR)
机构 * The University of Melbourne(墨尔本大学) ; The University of Western Australia(西澳大学)
专题命中 图文多模态 :multimodal(title);分类 cs.CL
Comments Findings of ACL 2025
机构 * University of Lincoln(林肯大学) ; University of Exeter(埃克塞特大学) ; AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments Medical Image Understanding and Analysis (MIUA) 2025 Extended Abstract Submission
机构 * Peng Wang and Shuai Bai and Sinan Tan and Shijie Wang and Zhihao Fan and Jinze Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Yang Fan and Kai Dang and Mengfei Du and Xuancheng Ren and Rui Men and Dayiheng Liu and Chang Zhou and Jingren Zhou and Junyang Lin(研究人员)
专题命中 图文多模态 :multimodal(title);分类 cs.CL
Comments 16 pages
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments 24 pages, 10 figures
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments This paper has been accepted for presentation at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments 19 pages, 13 figures, 3 tables
专题命中 图文多模态 :multi-modal(title);分类 cs.AI
Comments 13 pages, 7 figures, 5 tables
Journal ref Health Data Science, 2024
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments Accepted by CVPR 2024
专题命中 图文多模态 :multimodal(title);分类 cs.CV
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments 9 pages, 6 figures, presented on NeurIPS workshop on Robustness of Few-shot and Zero-shot Learning in Foundation Models
专题命中 图文多模态 :multimodal(title);分类 cs.AI
专题命中 图文多模态 :image-text(title);分类 cs.CV
Comments CVPR 2023
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
专题命中 图文多模态 :image-text(title);分类 cs.CV
Comments 8 pages, 8 figures, Accepted at SIGGRAPH ASIA 2022, Project Page at https://www.nasir.lol/clipmesh
专题命中 图文多模态 :image-text(title);分类 cs.CV
Comments 10 pages, 4 figures, presented in the MACLEAN workshop during ECML PKDD 2021
专题命中 图文多模态 :multi-modal(title);分类 cs.CL
Comments 8 pages, 2 figures, 2 tables, accepted to NAACL-HLT SRW 2021
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
专题命中 图文多模态 :image-text(title);分类 cs.MM
Comments Accepted by ACMMM2019
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments ICCV 2019
专题命中 图文多模态 :multi-modal(title);分类 cs.CV
Comments 6 pages, accepted at MIPR 2019
专题命中 图文多模态 :multimodal(title);分类 cs.CV
Comments 4 pages, 1 figure, accepted at MIPR2018
专题命中 图文多模态 :image-text(title);分类 cs.CL
StateSight:评测视觉语言模型中的潜在空间状态重建能力
机构 * Thomas Jefferson High School for Science and Technology(托马斯·杰斐逊科学技术高中)
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 本研究推出StateSight基准及配套数据集StateSight-Steps,评估视觉语言模型的潜在空间状态重建能力,发现GPT-5.5、Claude Sonnet 5的表现均逊于人类基线,格式正确的响应可能掩盖空间结构恢复失败的问题。
回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉
机构 * Xidian University(西安电子科技大学) ; Tsinghua University(清华大学) ; Shanghai Road Transport Development Center(上海市道路运输发展中心) ; Hunan Institute of Advanced Technology(湖南先进技术研究院)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM
AI总结 研究针对LVLMs易产生幻觉问题,提出无需训练的SDPR框架,通过显著性驱动注意力重新分配、缓存对齐及先验约束对比解码,整体对齐视觉意识,在多基准测试中优于现有方法,无需额外训练且开销小。
Comments Accepted by ACM Multimedia 2026