arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SignDino:通过时间轴自蒸馏实现自监督手语表示学习

SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

Junyi Hu, Zhewen He, Haomian Huang, Zhenhua Li, Zhifei Li, Yi Fang

arXiv 2609.06296首次发表:更新:

发表机构

New York University Abu Dhabi(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SignDino通过时间轴自蒸馏将DINOv3范式迁移至手语视频,分解手、脸流并训练轻量时间Transformer,在翻译、识别和手指拼写任务上达到先进性能。

AI 中文摘要

自监督手语表示学习必须建模自然图像自监督学习不关注的两个特性:手语由一组数量有限、解剖学上不同的发音器官产生,且其含义取决于这些发音器官的时间组织。我们提出了SignDino,一种自监督手语视频编码器,它将DINOv3的学生-教师范式从图像裁剪的空间域迁移到跟踪手语流的时间域。每个视频通过一个检测器优先的YOLOv8n+ByteTrack流程分解为左手、右手和面部流。一个冻结的DINOv3 ViT-B/16嵌入每个逐帧解剖裁剪,而轻量级的时间Transformer(而非图像骨干网络)构成学生和EMA教师。它们通过时间DINO自蒸馏、iBOT风格的帧级掩码标记预测、KoLeo特征扩展以及帧间相似性结构的Gram锚定进行训练。这种设计保持强图像级视觉基元固定,仅学习发音器官状态如何随时间演变。我们在手语到英语翻译、孤立手语识别和手指拼写检测基准上进行了评估。在这些任务中,SignDino提供了强大的公开自监督表示,并在匹配的下游评估下展现出具有竞争力或最先进的性能。

英文摘要

Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.

Comments24 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑