Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
视觉语言动作模型所言即所指?论忠实性在具身推理中的作用
Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone
机构
*
Stanford University(斯坦福大学)
;
Politecnico di Milano(米兰理工大学)
;
KTH Royal Institute of Technology(皇家理工学院)
;
Italian Institute of Artificial Intelligence (AI4I)(意大利人工智能研究所)
;
NVIDIA Research(英伟达研究院)
How to Build Digital Humans? From Priors to Photorealistic Avatars
如何构建数字人?从先验知识到逼真的虚拟形象
Wojciech Zielonka, Tobias Kirschstein, Timo Bolkart, Simon Giebenhain, Vanessa Sklyarova, Xiang Deng, Donglai Xiang, Shunsuke Saito, Yebin Liu, Matthias Niessner, Justus Thies
机构
*
Meta
;
Technical University of Munich(慕尼黑技术大学)
;
Google(谷歌)
;
Tsinghua University(清华大学)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
NVIDIA
;
Technical University of Darmstadt(达姆施塔特技术大学)
;
ETH Zurich(苏黎世联邦理工学院)
机构
*
Nanjing University(南京大学)
;
Australian National University(澳大利亚国立大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Nvidia(英伟达)