Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
机构 * POSTECH ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments EMNLP 2025 Main Conference
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * POSTECH ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments EMNLP 2025 Main Conference
机构 * Netflix, Inc.(Netflix公司) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV
Comments 11 pages, 5 figures, 5 tables
机构 * Shenzhen University(深圳大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465
专题命中 视觉推理 :visual reasoning(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments ICCV 2025 MARS2 Workshop and Challenge "Multimodal Reasoning and Slow Thinking in the Large Model Era: Towards System 2 and Beyond''
机构 * University of Naples Federico II(那不勒斯费德里科二世大学) ; Northwestern University(西北大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
Comments Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025
机构 * Zhejiang University(浙江大学)
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) ; College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) ; CMIC, Shanghai Jiao Tong University(上海交通大学CMIC) ; Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心) ; Center for Future Media, University of Electronic Science and Technology of China(电子科技大学未来媒体中心)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)
专题命中 视觉定位与Grounding :grounding(title);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Technical Report. Project Page: https://klingavatar.github.io/
机构 * Department of Electronic Systems, Aalborg University, Denmark(电子系统系,奥胡斯大学) ; Department of Health Science and Technology, Aalborg University, Denmark(健康科学与技术系,奥胡斯大学)
专题命中 视觉定位与Grounding :visual language model(title);vision language model(abstract);VLM(abstract)
Comments ICAT 2025
机构 * Bilibili(哔哩哔哩) ; UESTC ; University of Virginia(弗吉尼亚大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by EMNLP2025 Finding
机构 * School of Automation, Southeast University(东南大学自动化学院) ; Baidu Inc.(百度公司) ; JiangNan University(江南大学) ; Nanjing University of Science and Technology(南京理工大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in September 2025
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI2025)
机构 * Zhejiang University(浙江大学)
专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.LG
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments 18 pages
机构 * University of Naples Federico II(那不勒斯费迪里奇二世大学) ; Northwestern University(西北大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Journal ref SIGIR '25: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025
专题命中 视觉定位与Grounding :vision language model(abstract)
机构 * University of Technology Sydney(悉尼技术大学) ; Tencent(腾讯) ; Beijing Jiaotong University(北京交通大学) ; Westlake University(西湖大学)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构 * UC San Diego(圣迭戈大学) ; Hillbot
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Project Page: https://gen-vla.github.io/
机构 * Peking University(北京大学) ; Galbot ; USTC(中国科学技术大学) ; BAAI(百度人工智能研究院) ; University of Adelaide(阿德莱德大学) ; Zhejiang University(浙江大学) ; Differential Robotics Project(差分机器人项目)
专题命中 GUI与屏幕智能体 :vision-language model(abstract)
Comments Project Page: https://pku-epic.github.io/NavFoM-Web/
机构 * School of Computer Engineering(计算机工程学院) ; Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) ; Efficient Computing, Xiamen University(高效计算,厦门大学) ; Tencent Youtu Lab(腾讯优图实验室)
专题命中 VLM训练与架构 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted by ICML 2025
机构 * Dept. of Electronic Engineering, Sogang University(电子工程系,成均馆大学) ; Dept. of Artificial Intelligence, Sogang University(人工智能系,成均馆大学)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract)
机构 * Shenzhen Technology University(深圳科技大学) ; University of Washington(华盛顿大学) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments 13 pages,12 figures
机构 * School of Software Engineering(软件工程学院) ; Guangdong Laboratory of Artificial Intelligence and Digital Economy(人工智能与数字经济广东实验室)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
机构 * Korea University(韩国大学)
专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS, Beijing, China(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院,北京,中国) ; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国) ; Fanyu AI Laboratory, Zhongke Fanyu Technology Co., Ltd, Beijing, China(凡语AI实验室,中科创始人技术有限公司,北京,中国)
专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV
Comments EMNLP2025 Main
机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) ; Stony Brook University(石溪大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
专题命中 其他VLM :MLLM(abstract)