Latent Expression Generation for Referring Image Segmentation and Grounding
机构 * GIST(韩国科学技术院) ; Seoul National University(首尔国立大学) ; POSTECH
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to ICCV 2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * GIST(韩国科学技术院) ; Seoul National University(首尔国立大学) ; POSTECH
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to ICCV 2025
机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) ; University of Adelaide(阿德莱德大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * Media Evaluation Lab, ByteDance Inc.(字节跳动公司媒体评估实验室)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
Comments Extended journal version of arXiv:2506.03662
机构 * Arizona State University(亚利桑那州立大学) ; Clemson University(克莱姆森大学) ; LinkedIn Corporation(领英公司) ; Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG
机构 * Keio University(庆应大学) ; University of Stuttgart(斯图加特大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments This work is a pre-print version of a paper that has been accepted to the IEEE International Symposium on Mixed and Augmented Reality for future publication. Project Page: https://mediated-reality.github.io/projects/yasunaga_ismar25/
机构 * Shenzhen Key Laboratory of Media Security, Faculty of Electronic and Information Engineering, Shenzhen University, China(深圳媒体安全重点实验室,电子与信息工程学院,深圳大学,中国) ; Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(快速富对象搜索(ROSE)实验室,电气电子工程学院,南洋理工大学,新加坡) ; Guangdong Laboratory of Machine Perception and Intelligent Computing, Faculty of Engineering, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,工程学院,深圳MSU-BIT大学,中国)
专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV
机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) ; University of California, Irvine(加州大学尔湾分校) ; University of California, Berkeley(加州大学伯克利分校) ; Covalent Metrology(协力计量)
专题命中 视觉定位与Grounding :visual reasoning(abstract);分类 cs.CV
Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop
专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV
机构 * Analytics Everywhere Lab, University of New Brunswick, Canada(新不伦瑞克大学分析 everywhere 实验室) ; University of Foreign Language Studies, University of Danang, Vietnam(越南丹绒大学外语学院) ; Faculty of Information Technology, University of Science, VNU-HCM, Vietnam(越南胡志明市大学信息科技学院) ; Faculty of Economics and Accounting, Quy Nhon University, Vietnam(越南奎隆大学经济与会计学院) ; Faculty of Natural Sciences, Quy Nhon University, Vietnam(越南奎隆大学自然科学学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Published as a conference paper at ICEFM 2025
机构 * Dept. of CSE PES University Bangalore, India(计算机科学与工程系,PES大学,印度班加罗尔)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract)
Comments This is a pre-peer-review version of a paper with the same title accepted at 20th IFIP TC13 International Conference on Human-Computer Interaction (INTERACT 2025)
专题命中 视觉定位与Grounding :grounding(abstract)
Comments Molecular Physics (2025)