arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无标注家具编码:它们编码了什么以及能转移多远

Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer

Benjamin Friedman

arXiv 2607.10461首次发表:更新:

发表机构

DLR Group(DLR集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究能否用源自物体几何形状的自监督令牌取代类别标签和规范姿态约定。通过对FSQ点云自动编码器进行倒角训练,发现编码能恢复多种信息,其旋转内容可由训练目标设置,跨资产库缩放时编码转移情况因类别而异。

AI 中文摘要

基于布局的3D场景合成器通过两个人工标注通道放置每个物体:类别标签和规范姿态约定。我们探讨是否一个源自物体几何形状的自监督令牌可以取代这两者,并将此类令牌作为一种独立于任何合成器的表示进行研究。使用无标签或姿态标注的放置好的3D-FUTURE家具对有限标量量化(FSQ)点云自动编码器进行倒角训练。诊断探针仅从编码中就能恢复精细类别(62.6 +/- 0.5%)、超类别(85.6 +/- 1.3%)和偏航(52.7 +/- 0.5度)。将倒角目标从旋转点云切换到未旋转点云会使偏航信号消失,同时提高类别恢复率,表明编码的旋转内容可由训练目标设置。跨资产库缩放需要可转移的编码;在一个未见数据集(ShapeNet)上,对齐取决于类别:盒状家具可转移,有机形状家具不可转移,一种目标盲增强部分缩小了差距。

英文摘要

Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose convention. We ask whether a single self-supervised token derived from object geometry can replace both, and study such tokens directly as a representation, decoupled from any synthesizer. A Finite Scalar Quantization (FSQ) point-cloud autoencoder is chamfer-trained on placed 3D-FUTURE furniture with no labels or pose annotations. Diagnostic probes recover fine-category (62.6 +/- 0.5%), super-category (85.6 +/- 1.3%), and yaw (52.7 +/- 0.5 deg) from the codes alone. Swapping the chamfer target from the rotated to the un-rotated point cloud collapses the yaw signal while raising class recovery, showing the codes' rotation content can be set by the training objective. Scaling across asset libraries needs codes that transfer; on an unseen dataset (ShapeNet), alignment is category-dependent: box-like furniture transfers, organically-shaped furniture does not, and a target-blind augmentation partly closes the gap.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑