机构
*
Multimodal Language Department(多模态语言部门)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
;
Department of Linguistics(语言学系)
;
Boğaziçi University(博多伊奇大学)
;
Donders Institute for Brain Cognition and Behaviour(多纳尔斯脑认知与行为研究所)
;
Radboud University(拉德堡德大学)
;
Department of Linguistics and Communication(语言学与沟通系)
;
University of Birmingham(伯明翰大学)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
从街景到视觉网络:利用视觉语言模型映射城市地标可见性
Zicheng Fan, Kunihiko Fujiwara, Pengyuan Liu, Fan Zhang, Filip Biljecki
机构
*
organization= Department of Architecture, National University of Singapore , country= Singapore
;
organization= Research \& Development Institute, Takenaka Corporation , country= Japan
;
organization= Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow , country= United Kingdom
;
organization= Institute of Remote Sensing
;
GIS, Peking University , country= China
;
organization= Department of Real Estate, National University of Singapore , country= Singapore
专题命中
视觉定位与Grounding
:vision-language model(title);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式
Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所)
;
PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
ENC-Bench:用于评估多模态大语言模型在电子航海图理解中的基准
Ao Cheng, Xingming Li, Xuanyu Ji, Xixiang He, Qiyao Sun, Chunping Qiu, Runke Huang, Qingyong Hu
机构
*
National University of Defense Technology(国防科技大学)
;
Intelligent Game and Decision Lab(智能游戏与决策实验室)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);InternVL(abstract);grounding(abstract);分类 cs.CV
EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
EagleVision:基于BEV的双阶段框架用于空间智能
Jiaxu Wan, Xu Wang, Mengwei Xie, Hang Zhang, Mu Xu, Yang Han, Hong Zhang, Ding Yuan, Yifan Yang
机构
*
Atlas Lab(Atlas实验室)
;
School of Aerospace, BUAA(北京航空航天大学航空学院)
;
School of Software, BUAA(北京航空航天大学软件学院)
;
State Key Laboratory of HERATT(HERATT国家重点实验室)
;
Key Laboratory of SDODS (MOE) Project(SDODS重点实验室(教育部))
专题命中
视觉定位与Grounding
:grounding(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
;
School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)
;
CSIRO(澳大利亚联邦科学与工业研究组织)
;
Northwestern Polytechnical University(西北工业大学)
;
University of Macau(澳门大学)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
基于多模态大语言模型的零样本人-物交互检测
Shiyu Xuan, Dongkai Wang, Zechao Li, Jinhui Tang
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院)
;
Nanjing Forestry University(南京林业大学)
机构
*
School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)
;
Low-Altitude Intelligence Lab, Xiong’an National Innovation Center(雄安国家创新中心低空智能实验室)
;
Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创莲田科技有限公司)
;
College of Electronic Science and Technology, National University of Defense Technology(国防科技大学电子科学学院)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV