TinyDETR-Pose:面向基于轻量Transformer的端到端实时单阶段6DoF物体位姿估计
TinyDETR-Pose: Towards End-to-End Real-Time Single-Stage 6DoF Object Pose Estimation with Lightweight Transformers
查看机构详情
- Fraunhofer IGD(弗劳恩霍夫计算机图形学研究所)
- TU Darmstadt(达姆施塔特工业大学)
- TU Delft(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出轻量Transformer框架TinyDETR-Pose,通过单阶段端到端方式实现6DoF物体位姿估计,在YCB-V数据集上取得85.9的ADD-S AUC,参数量较同类方法减少72.7%,在Jetson Nano上每帧推理延迟仅4.5毫秒,可满足边缘设备实时部署需求。
中文摘要 AI 辅助
在资源受限硬件上实现实时6DoF物体位姿估计仍是一项挑战,因为基于准确对应关系的优化管线通常依赖不可微的PnP/RANSAC阶段或代价高昂的迭代优化,而近期基于基础模型的方法会产生难以在边缘设备部署的推理成本。本文提出TinyDETR-Pose,这是一个轻量、端到端的单阶段框架,可在单次前向传播中同时完成物体检测和完整6DoF位姿回归。该框架构建于高效的LW-DETR架构之上,将检测和位姿估计建模为集合预测问题,并为每个解码器查询附加专用MLP头,用于旋转、单目深度及投影物体中心的回归,无需使用PnP、NMS(非极大值抑制)或迭代位姿优化。物体对称性通过对所有物体统一应用的ADD-S损失处理,无需针对特定物体设置单独的损失调度或分开的测地线/ADD监督。此外,预测结果通过基于类别和2D空间线索的对称性安全匈牙利匹配器分配给真实值,在对称性和深度模糊情况下仍能实现稳定分配。在YCB-V数据集上,TinyDETR-Pose取得了85.9的可比ADD-S AUC,同时比其他基于DETR的单阶段位姿估计方法减少了多达72.7%的参数。由于其紧凑的设计,TinyDETR-Pose可实现实时运行,在NVIDIA Jetson Nano上使用TensorRT时,每帧推理延迟仅约4.5毫秒,证明基于Transformer的准确端到端6DoF位姿估计可实际应用于边缘部署。
英文摘要
Real-time 6DoF object pose estimation on resource-constrained hardware remains challenging, as accurate correspondence-based and refinement pipelines typically rely on non-differentiable PnP/RANSAC stages or costly iterative refinement, while recent foundation-model-based approaches incur inference costs that are prohibitive for edge deployment. We present TinyDETR-Pose, a lightweight, end-to-end, single-stage framework that jointly detects objects and regresses their full 6D pose in a single forward pass. Built on the efficient LW-DETR architecture, TinyDETR-Pose formulates detection and pose estimation as a set-prediction problem and attaches dedicated MLP heads for rotation, monocular depth, and projected object center regression to each decoder query, eliminating the need for PnP, NMS (non-maximum suppression), or iterative pose refinement. Object symmetries are handled through a ADD-S loss applied uniformly to all objects, without the need for object-specific loss schedules or separate geodesic/ADD supervision. In addition, predictions are assigned to ground truth using a symmetry-safe Hungarian matcher based on class and 2D spatial cues, yielding stable assignment under symmetry and depth ambiguity. On YCB-V, TinyDETR-Pose achieves a comparable ADD-S AUC of 85.9, while requiring up to 72.7% fewer parameters than other DETR-based single-stage pose-estimation approaches. Due to its compact design, TinyDETR-Pose runs in real time and achieves an inference latency of only ~4.5 ms per frame on an NVIDIA Jetson Nano using TensorRT, demonstrating that accurate end-to-end transformer-based 6D pose estimation can be made practical for edge deployment.