arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一统所有:面向跨传感器骨架表示学习的通用基础模型

One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning

Jeonghyeok Do, Yun Chen, Munchurl Kim

arXiv 2609.07078首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院(KAIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对骨架数据异质性导致需训练多个模型的问题,提出通用基础模型SOfA,利用规范关节槽和语义关节嵌入统一跨传感器表示,在十个数据集上以单一模型实现SOTA性能。

AI 中文摘要

为了从大规模无标注数据中学习可泛化的运动表示,自监督学习(SSL)已成为一种广泛采用的方法。然而,现有方法主要受限于骨架数据固有的异质性——其特点是不同传感器下关节数量、索引协议和拓扑结构各不相同——这通常需要训练独立的、特定于传感器的甚至完全特定于数据集的模型。为了克服这一问题,我们引入了SOfA(Skeleton One for All,骨架一统所有),这是首个旨在实现跨不同传感器统一骨架表示学习的通用基础模型。为了适应由不同关节数量引起的维度差异,我们引入了一组固定大小的可学习规范关节槽(Canonical Joint Slots),作为通用容器,无缝容纳任意骨架拓扑。SOfA通过注意力机制填充这些槽,动态聚合来自特定传感器输入的骨架信息。此外,我们通过引入源自预训练文本编码器的语义关节嵌入(Semantic Joint Embedding)来解决不同传感器之间关节索引不对齐的问题,而非依赖绝对位置嵌入。为验证我们的方法,我们标准化了十个3D骨架数据集进行统一训练。大量实验表明,SOfA可充当真正通用的编码器,在广泛的下游任务和传感器类型上实现最先进(SOTA)性能,且通常以单一基础模型超越特定数据集的专业模型。

英文摘要

For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (SSL) has become a widely adopted methodology. However, existing approaches are primarily limited by the inherent heterogeneity of skeleton data---characterized by varying joint counts, indexing protocols, and topological structures across different sensors---which typically necessitates training separate, sensor-specific, or even entirely dataset-specific models. To overcome this, we introduce SOfA (Skeleton One for All), the first generalist foundation model designed to achieve sensor-unified skeleton representation learning across diverse sensors. To accommodate the dimensional gap caused by varying joint counts, we introduce a fixed-size set of learnable Canonical Joint Slots, acting as a universal vessel that seamlessly accommodates arbitrary skeletal topologies. SOfA fills these slots via an attention mechanism that dynamically aggregates skeletal information from sensor-specific inputs. Furthermore, we resolve joint index misalignment between various sensors by introducing a Semantic Joint Embedding derived from a pre-trained text encoder, rather than relying on absolute positional embeddings. To validate our approach, we standardized ten 3D skeleton datasets for unified training. Extensive experiments demonstrate that SOfA can serve as a truly universal encoder, achieving state-of-the-art (SOTA) performance across a wide range of downstream tasks and sensor types, often outperforming dataset-specific specialist models with a single foundation model.

CommentsPlease visit our project page at https://kaist-viclab.github.io/SOfA_site/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑