PQR3D:多视角3D目标检测中基于参考条件时间窗口的渐进式查询细化
PQR3D: Progressive Query Refinement over Reference-Conditioned Temporal Windows for Multi-View 3D Object Detection
AI总结:
PQR3D通过参考条件时间窗口内的渐进式查询细化,实现无需持久记忆的随机帧采样和独立推理,在nuScenes上以ViT-L达到71.6 NDS和64.9 mAP的新纪录。
AI中文摘要:
时间上下文对于纯相机多视角3D目标检测至关重要。现有的流式检测器维护并传播从一帧到下一帧的查询状态,这要求序列感知训练和按时间顺序推理。我们提出了PQR3D,它在参考条件时间窗口内执行渐进式查询细化。这种设计支持随机帧采样和独立推理,无需持久查询记忆。在每个窗口内,PQR3D将运动对齐的高置信度查询从较早时间戳逐步转移到目标帧。我们进一步引入掩码自注意力来调节常规、传播和去噪查询之间的交互,同时保持去噪监督与检测查询隔离。此外,阶段解耦的锚点嵌入在自注意力之前注入位置,之后注入大小、方向和速度,从而减少时间不一致属性的干扰。使用ViT-L骨干网络,PQR3D在nuScenes测试集上达到了新的最先进水平,实现了71.6 NDS和64.9 mAP。源代码可在该https URL获取。
英文摘要:
Temporal context is essential for camera-only multi-view 3D object detection. Existing streaming detectors maintain and propagate query states from one frame to the next, requiring sequence-aware training and chronological inference. We propose PQR3D, which performs progressive query refinement within referenceconditioned temporal windows. This design enables random frame sampling and independent inference without persistent query memory. Within each window, PQR3D progressively transfers motion-aligned high-confidence queries from earlier timestamps toward the target frame. We further introduce masked selfattention to regulate interactions among regular, propagated, and denoising queries while keeping denoising supervision isolated from detection queries. In addition, a stage-decoupled anchor embedding injects position before self-attention and size, orientation, and velocity afterward, reducing interference from temporally inconsistent attributes. With a ViT-L backbone, PQR3D sets a new state of the art on the nuScenes test set, achieving 71.6 NDS and 64.9 mAP. Source code is available at https://github.com/huiyegit/PQR3D