HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning
HyperVis:洛伦兹双曲面上的连续潜在视觉关系图用于组合推理
Moshiur Farazi, Sameera Ramasinghe, Mahbub Ahmed Turza, Shafin Rahman
机构
*
Data Science and AI, University of Doha for Science and Technology, Qatar(数据科学与人工智能,多哈科学技术大学,卡塔尔)
;
Pluralis Research, Australia(Pluralis研究,澳大利亚)
;
Department of Electrical and Computer Engineering, North South University, Bangladesh(电气与计算机工程系,北南大学,孟加拉国)
CommentsWithdrawn by the authors due to pending intellectual property considerations. The authors have determined that the current version contains material that should not have been publicly disseminated at this stage
机构
*
Arizona State University(亚利桑那州立大学)
;
Clemson University(克莱姆森大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Halmstad University(哈姆斯塔德大学)
;
Florida State University(佛罗里达州立大学)
;
Rice University(里士满大学)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV
Rethinking Video-Language Model from the Language Input Perspective
从语言输入角度重新思考视频-语言模型
Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Wuhan University(武汉大学)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
如何以及想象什么?统一多模态模型中的视觉思维用于跨视角空间推理
Qian Yang, Ankur Sikarwar, Huy Le, Le Zhang, Zhuan Shi, Perouz Taslakian, Aishwarya Agrawal
机构
*
Mila - Québec AI Institute(蒙特利尔AI研究所)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
ServiceNow AI Research(ServiceNow人工智能研究)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
GeoSolver: 利用细粒度过程监督扩展遥感中的测试时推理
Lang Sun, Ronghao Fu, Zhuoran Duan, Haoran Liu, Xueyan Liu, Bo Yang
机构
*
College of Computer Science and Technology(计算机科学与技术学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University(教育部符号计算与知识工程重点实验室)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
Tsinghua University, SIGS(清华大学 SIGS)
;
Meituan(美团)
;
The Chinese University of Hong Kong(香港中文大学)
;
National University of Singapore(新加坡国立大学)
;
LMMs-Lab(LMMs实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
机构
*
Computer Vision Institute, School of Computer Science and Software Engineering, Shenzhen University(计算机视觉研究院,计算机科学与软件工程学院,深圳大学)
;
Guangdong Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室,深圳大学)
;
School of Computer Science, University of Nottingham Ningbo China(Nottingham Ningbo 中国计算机科学学院)
;
Department of Electrical and Computer Engineering, National University of Singapore(电子与计算机工程系,新加坡国立大学)
;
Department of Radiation Oncology, Stanford University(放射肿瘤科,斯坦福大学)
;
Sun Yat-sen University(中山大学)
;
School of Computer Science, University of Nottingham(计算机科学学院,Nottingham大学)