arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AQ3D:用于3D实例分割的自适应查询Transformer

AQ3D: Adaptive Query Transformer for 3D Instance Segmentation

Keno Moenck, Thorsten Schüppstuhl

arXiv 2608.30618首次发表:更新:

发表机构

Hamburg University of Technology(汉堡工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AQ3D,通过自适应查询和改进的位置编码、解码器模块,在无额外数据增强的情况下,于多个ScanNet数据集的3D实例分割任务中取得新的SOTA性能。

AI 中文摘要

基于Transformer的3D实例分割解码器通常采用固定数量的查询,且位置建模基于训练分布而非当前场景。室内扫描的空间范围和物体数量差异极大,固定查询集会导致小场景初始化过度、大场景初始化不足,而学习到的绝对和相对编码受限于训练场景范围,易出现饱和。本文提出AQ3D,可在训练和推理阶段处理不同规模的场景:查询按场景超点的固定比例实例化,形成过完备集合,背景剔除完全由解码器负责;位置信息采用量化度量坐标上的3D RoPE编码,替代以往解码器中学习到的有界查找表;进一步通过基于归因的超点池化、掩码细化分支及余弦分类器改进解码器以实现背景剔除。实验表明,在未使用额外数据增强训练的解码器方法中,该方法在ScanNetV2、ScanNet200和ScanNet++V2数据集的验证集和隐藏测试集上达到新的SOTA。代码可访问该网址获取。

英文摘要

Transformer-based decoders for 3D instance segmentation typically commit to a fixed number of queries and positional modeling calibrated on the training distribution rather than on the scene at hand. Indoor scans vary widely in spatial extent and object count, so a fixed query set over-initializes small scenes and under-initializes large ones, while learned absolute and relative encodings are bound to the training scenes' extents and can saturate. We present AQ3D, which is designed to handle scenes of various sizes during training and inference. Queries are instantiated at a fixed ratio of the scene's superpoints, forming an overcomplete set whose background rejection is entirely left to the decoder. Positional information is encoded using 3D RoPE over quantized metric coordinates, replacing learned bounded lookup tables of prior decoders. Further, we improve the decoder itself by using attribution-based superpoint pooling, a mask refinement branch, and a cosine classifier for background rejection. Experiments show our method sets a new state-of-the-art on validation and hidden test splits across the datasets ScanNetV2, ScanNet200, and ScanNet++V2 among decoder methods trained without additional data augmentation. Code is available at \href{https://github.com/kenomo/aq3d}{github.com/kenomo/aq3d}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑