Visual Jigsaw Post-Training Improves MLLMs
机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; Linköping University(_linköping大学) ; SenseTime Research(商汤科技研究院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; Linköping University(_linköping大学) ; SenseTime Research(商汤科技研究院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * New York University(纽约大学) ; NYU Shanghai(纽约大学上海分校) ; The University of Hong Kong(香港大学) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments 25 pages, 5 figures
机构 * The University of Tokyo(东京大学)
专题命中 视觉推理 :VLM(abstract);分类 cs.CV
Comments This paper has been accepted to ICCV 2025
机构 * University of Trento(特伦托大学) ; BIFOLD and Technische Universität Berlin(BIFOLD和柏林技术大学) ; Technical University of Munich(慕尼黑技术大学) ; University of Pisa(比萨大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * Shenzhen Key Laboratory of Robotics Perception and Intelligence(机器人感知与智能深圳重点实验室) ; Department of Electronic and Electrical Engineering(电子与电气工程系) ; Southern University of Science and Technology(南方科技大学) ; Research Institute of Multiple Agents and Embodied Intelligence(多智能体与具身智能研究院) ; Peng Cheng Laboratory(鹏城实验室) ; Robotics and Microsystems Center(机器人与微系统中心) ; School of Mechanical and Electric Engineering(机械与电子工程学院) ; Soochow University(苏州大学) ; Xiamen Key Laboratory of Visual Perception Technology and Application(视觉感知技术与应用厦门重点实验室) ; Reconova Information Technology Co., Ltd.(Reconova信息技术有限公司)
专题命中 视觉推理 :vision-language model(abstract)
专题命中 视觉推理 :grounding(abstract)
机构 * Institute for Infocomm Research (I 2 {}^{\text{2}} R), Agency for Science, Technology and Research (A*STAR)(信息与通信研究机构(I2R),科技研究局(A*STAR)) ; Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(前沿人工智能研究中心(CFAR),科技研究局(A*STAR))
专题命中 视觉推理 :grounding(abstract)
机构 * Southwest Jiaotong University(西南交通大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Tongji University(同济大学)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) ; Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) ; Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)
专题命中 视觉推理 :multimodal large language model(abstract)
Comments EMNLP 2025 Findings
机构 * Department of Computer Vision, Mohamed Bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed Bin Zayed人工智能大学) ; School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) ; Bytedance Seed(字节跳动种子) ; School of Computer Science, Qinghai Normal University(计算机科学学院,青海师范大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.AI
机构 * University of California, San Diego(加州大学圣地亚哥分校) ; ByteDance(字节跳动) ; University of California, Merced(加州大学默塞德分校) ; University of Southern California(南加州大学) ; University at Buffalo(布法罗大学) ; The University of Queensland(昆士兰大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * Tsinghua University(清华大学) ; Northwestern Polytechnical University(西北工业大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究所) ; China Telecom(中国电信) ; Zhejiang University(浙江大学) ; Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG
Comments accepted by ACL'2025
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Project page: https://seung-hun-lee.github.io/projects/TGL/
机构 * Sun Yat-sen University(中山大学) ; Peng Cheng Laboratory(鹏城实验室) ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Key Laboratory of Machine Intelligence and Advanced Computing(人工智能与先进计算重点实验室)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV
Comments 12 pages, 4 figures, Accepted by ICCV2025
机构 * Sun Yat-sen University(中山大学) ; School of Computing and Data Science(计算与数据科学学院) ; The University of Hong Kong(香港大学)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted at NeurIPS 2025 (preview; camera-ready in preparation)
机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术重点实验室,自动化研究所,中国科学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学) ; State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(认知神经科学与学习国家重点实验室,北京师范大学)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments Preprint submitted to Elsevier
机构 * AIR, Tsinghua University(清华大学) ; Peking University(北京大学) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.AI
机构 * University of Mainz(马尔堡大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.AI
Comments Paper presented at AMTA 2025
机构 * University of Padova(帕多瓦大学) ; Stanford University(斯坦福大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Pennsylvania(宾夕法尼亚大学) ; Örebro University(奥雷布罗大学) ; NVIDIA Research(NVIDIA研究)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)
Comments Project page: https://j-dapt.github.io/. 9 pages, 6 figures
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)
Comments https://github.com/UBC-NLP/pearl
机构 * Korea University(韩国大学) ; University of Seoul(首尔大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments 23 pages, 17 figures
机构 * Fractal AI Research(Fractal AI研究)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
机构 * KAIST(韩国科学技术院) ; Seoul National University(首尔国立大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments Accepted to EMNLP 2025 (Main Conference)
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) ; Beijing Electronic Science and Technology Institute(北京电子科学技术研究所) ; Wangxuan Institute of Computer Technology(北京大学计算机技术研究院) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Nanjing University of Science and Technology(南京理工大学) ; Beijing Normal University(北京师范大学)
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI
Comments 10 pages, 5 figures
机构 * University of Maryland, College Park(马里兰大学学院公园分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG
机构 * Wuhan University of Technology(武汉理工大学) ; Northwestern Polytechnical University(西北工业大学) ; Tsinghua University(清华大学)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai AI Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
机构 * Faculty of Business and Information Technology(商业与信息技术学院) ; Ontario Tech University(安大略技术大学)
专题命中 视觉定位与Grounding :grounding(abstract)
Comments Accepted for presentation at the LLMs Meet Databases (LMD) Workshop, 35th IEEE International Conference on Collaborative Advances in Software and Computing, 2025. Workshop website: https://sites.google.com/view/lmd2025/home