KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
KG-ViP:在多模态大语言模型中桥接知识基础与视觉感知以进行视觉问答
Zhiyang Li, Ao Ke, Yukun Cao, Xike Xie
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Data Darkness Lab, MIRACLE Center, USTC(数据黑暗实验室,MIRACLE中心,中国科学技术大学)
;
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
DinoComplete: 3D Shape Completion with Distilled Semantic Priors and State Space Models
DinoComplete: 利用蒸馏语义先验和状态空间模型进行3D形状补全
Furkan Mert Algan, Eckehard Steinbach
机构
*
Chair of Media Technology(媒体技术教授职位)
;
Munich Institute of Robotics and Machine Intelligence(慕尼黑机器人与机器智能研究所)
;
School of Computation Information and Technology, Technical University of Munich(计算信息科学学院,慕尼黑技术大学)
CommentsThis manuscript has been withdrawn by the authors because we found a methodological flaw in the formulation and evaluation of the proposed approach. The issue affects the reliability of the experimental results and the conclusions drawn from them. Therefore, the authors consider the current version unsuitable for citation or further use
Language Movement Primitives: Grounding Language Models in Robot Motion
语言运动基元:将语言模型锚定在机器人运动中
Yinlong Dai, Benjamin A. Christie, Daniel J. Evans, Dylan P. Losey, Simon Stepputtis
机构
*
Collab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(合作组,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
;
TEA Lab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(TEA实验室,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
机构
*
Tsinghua University, SIGS(清华大学 SIGS)
;
Meituan(美团)
;
The Chinese University of Hong Kong(香港中文大学)
;
National University of Singapore(新加坡国立大学)
;
LMMs-Lab(LMMs实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
机构
*
Zhejiang Key Laboratory of Space Information Sensing and Transmission(浙江空间信息感知与传输重点实验室)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
Zhejiang University(浙江大学)
;
Tsinghua University(清华大学)
;
Children's Hospital, Zhejiang University School of Medicine(浙江大学医学院附属儿童医院)
3D LULC classification using multispectral LiDAR and deep learning: current and prospective schemes
基于多光谱LiDAR和深度学习的3D土地利用/覆盖分类:当前和未来方案
Narges Takhtkeshha, Aldino Rizaldy, Markus Hollaus, Juha Hyyppä, Fabio Remondino, Gottfried Mandlburger
机构
*
D Optical Metrology (3DOM) Unit, Bruno Kessler Foundation (FBK)(3D光学计量(3DOM)单元,布鲁诺·凯塞尔基金会(FBK))
;
Department of Geodesy and Geoinformation, TU Wien(测绘与地理信息系,维也纳技术大学)
;
Helmholtz-ZentrumDresden-Rossendorf (HZDR), Helmholtz Institute Freiberg for Resource Technology (HIF)(德累斯顿-罗斯托克赫尔姆霍尔茨研究中心(HZDR),资源技术赫尔姆霍尔茨研究所(HIF))
;
Freie Universität Berlin, Remote Sensing and Geoinformatics(柏林自由大学,遥感与地理信息学)
;
Department of Remote Sensing and Photogrammetry, Finnish Geospatial Research Institute FGI, The National Land Survey of Finland(芬兰地理研究所(FGI),芬兰国家土地测绘局)
One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems
一句话,一出戏剧:通过多智能体系统实现个性化短剧生成
Yufei Shi, Weilong Yan, Naixuan Huang, Yucheng Chen, Chenyu Zhang, Tao He, Si Yong Yeo, Ming Li
机构
*
MedVisAI Lab, Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室,李光前医学院,南洋理工大学)
;
National University of Singapore(新加坡国立大学)
;
Beijing Institute of Technology(北京理工大学)
;
Tsinghua University(清华大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Guangming Laboratory(光明实验室)