探究预训练检测Transformer的三维物体级理解能力
Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
浏览论文内容
中文总结 AI 辅助
本文探究预训练二维检测Transformer对物体三维属性的理解,发现无三维监督预训练的二维DETR模型,能通过线性和非线性探针从物体级嵌入中恢复物体深度、三维位置等三维属性信息,具备强的三维物体级理解能力。
中文摘要 AI 辅助
包括DETR及其扩展在内的检测Transformer模型,学习输出一组物体级嵌入,这些嵌入可同时解码为二维边界框和类别分布。本文研究预训练二维检测Transformer对物体三维属性的理解,具体探究从物体级嵌入中使用线性和非线性探针可恢复的属性范围,包括物体到相机的深度、物体相对于相机的三维位置。在一系列检测Transformer模型上的实验结果显示,尽管模型预训练过程中完全没有三维监督,二维DETR模型仍具备此前未知的、强得惊人的表示物体三维属性有用信息的能力。
英文摘要
Detection transformer models, including DETR and its extensions, learn to output a set of object-level embeddings that can be simultaneously decoded into 2D bounding boxes and class distributions. In this paper, we investigate what pre-trained 2D detection transformers understand about the 3D properties of objects. Specifically, we investigate the extent to which properties including the depth of objects from the camera and the 3D location of objects relative to the camera can be recovered from object-level embeddings using linear and non-linear probes. Across a range of detection transformer models, our results show a surprisingly strong and previously unknown ability of 2D DETR models to represent useful information about the 3D properties of objects, despite the complete lack of 3D supervision during model pre-training.
发表机构
- UMass Amherst(马萨诸塞大学阿默斯特分校)
- SRI International(SRI国际公司)
机构由 AI 辅助整理,请以论文原文为准。