arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无 gloss( gloss 指手语 gloss,即手语的书面转写)的跨数据集手语定位表征学习

Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

Oğuz Akif Tüfekcioğlu, Ezgi Ekin, Mustafa Kaan Çevik, Hacer Yalim Keles

arXiv 2608.11332首次发表:更新:

发表机构

Hacettepe University(哈杰泰佩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对资源受限语言的手语研究,用土耳其广播语料库 TSL-News 基于转录文本的伪 gloss 标签预训练手语编码器,在跨数据集手语定位任务中显著提升性能,证明松散对齐的广播数据可提供有效弱监督学习手语表征。

AI 中文摘要

针对资源受限语言的手语研究常受限于密集语言标签(如 gloss、时间边界、手语顺序)的标注成本。广播新闻提供了一种实用替代方案,其将连续手语与口语 transcript(转录文本)配对,但由于文本与手语的对齐较为松散,这种监督信号较弱。土耳其等形态丰富的语言进一步增加了难度:同一词汇含义可呈现为多种屈折形式,而某些派生形式需保持区分。本研究探讨在此场景下,基于弱转录监督能否预训练出可复用的手语编码器,其中较差的文本归一化会碎片化伪 gloss 目标,削弱表征学习效果。与以往主要用于改进翻译的伪 gloss 流水线不同,本研究测试预训练编码器能否作为可复用表征迁移至跨数据集手语定位任务。我们在新的土耳其广播语料库 TSL-News 上进行预训练,使用从转录文本衍生的伪 gloss 标签而非人工标注,对比基于规则的形态词形还原与基于固定词汇的约束大语言模型(LLM)辅助归一化两种方法。我们通过在由 TSL Dictionary 语料库构建的新 TSL Spotting Benchmark 上开展跨数据集手语定位,评估学习到的表征。结果显示,LLM 辅助编码器将 top-5 时间定位平均交并比(mean IoU)从 0.235 提升至 0.465,56.2% 的样本达到至少 0.50 的 IoU;频率分析表明,该提升并非主要由记忆高频伪 gloss 标签驱动。在下游翻译检验中,相同预训练将 BLEU-4 从 9.60 提升至 11.04,ROUGE 从 23.48 提升至 27.43。这些结果表明,对齐松散的广播数据可提供有效弱监督,用于学习兼具词汇内容与时间结构的手语表征。

英文摘要

Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak since text and signing are loosely aligned. Morphologically rich languages such as Turkish add further difficulty, as the same lexical meaning can appear in many inflected forms while some derived forms should remain distinct. We study whether weak transcript-based supervision can pretrain a reusable sign encoder in this setting, where poor text normalization can fragment pseudo-gloss targets and weaken representation learning. Unlike prior pseudo-gloss pipelines designed mainly to improve translation, we test whether the pretrained encoder transfers as a reusable representation for cross-dataset sign spotting. We pretrain on TSL-News, a new Turkish broadcast corpus, using pseudo-gloss labels derived from transcripts rather than manual annotation, comparing rule-based morphological lemmatization with constrained LLM-assisted normalization over a fixed vocabulary. We evaluate the learned representations via cross-dataset sign spotting on a new TSL Spotting Benchmark built from the TSL Dictionary corpus. The LLM-assisted encoder raises top-5 temporal localization mean IoU from 0.235 to 0.465, with 56.2% of examples reaching an IoU of at least 0.50; a frequency analysis suggests this gain is not mainly driven by memorizing frequent pseudo-gloss labels. In a downstream translation check, the same pretraining improves BLEU-4 from 9.60 to 11.04 and ROUGE from 23.48 to 27.43. These results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure.

CommentsAccepted at the 4th LIMIT Workshop (Representation Learning with Very Limited Resources), ECCV 2026. The abstract was shortened to comply with arXiv's 1,920-character limit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑