The Telephone Game: Evaluating Semantic Drift in Unified Models
机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学)
专题命中 VLM训练与架构 :VLM(abstract);visual language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学)
专题命中 VLM训练与架构 :VLM(abstract);visual language model(abstract);分类 cs.CV
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Harvard University(哈佛大学) ; Johns Hopkins University(约翰霍普金斯大学) ; Brown University(布朗大学)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.AI
Comments 8 pages, 7 figures
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(多模态人工智能系统国家重点实验室(MAIS),中国科学院自动化所) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院AutoLab) ; Anyverse Intelligence ; Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京多模态信息超智能安全重点实验室) ; School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV
Comments 13 pages, 2 figures
Journal ref NeurIPS 2025
机构 * Stanford University(斯坦福大学) ; UC Berkeley(加州大学伯克利分校) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments The code is available at https://github.com/bronyayang/Law_of_Vision_Representation_in_MLLMs
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.LG
机构 * School of Electronic Engineering(电子工程学院) ; Xidian University(西安电子科技大学) ; School of Computer Science(计算机科学学院)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
机构 * Sichuan University(四川大学) ; Peking University(北京大学)
专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract)
机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) ; Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security
专题命中 其他VLM :multimodal large language model(abstract)
Comments ACM MM 2025