arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06843cs.CV

RAIDAL:基于CTC的连续手语识别中的冗余感知信息密度主动学习

RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition

  • Universidade Federal de Minas Gerais(米纳斯吉拉斯联邦大学)
  • University of Bristol(布里斯托大学)

机构由 AI 辅助整理,请以论文原文为准。

Rafael A. Diniz Augusto, Gabriel L. Oliveira, Erickson R. Nascimento

AI总结:

针对连续手语识别中视频标注成本高的问题,提出RAIDAL方法,利用CTC解码器对齐峰值限制评分区域,在三个数据集上显著提升数据效率。

AI中文摘要:

连续手语识别(CSLR)是无障碍访问的关键技术,但其发展仍受到连续视频流标注成本高昂的限制。主动学习为降低这一成本提供了一条途径,但标准的采集函数并非为弱对齐的手语视频而设计,在手语视频中,手语动作与休息姿态、不规则停顿、类似手语的动作以及时间上冗余的帧交错出现。这种时间冗余可能削弱样本选择的效果,因为采集分数可能受到与解码出的词汇序列无关区域的时间步影响,从而扭曲视频估计的信息量。在本工作中,我们表明现代CSLR模型已经包含了一种识别词汇级时间证据的机制:CTC解码器。尽管CTC解码器通常仅在推理时使用,但其对齐峰值指示了模型在特征序列中定位每个预测词汇的位置,为主动学习采集函数提供了时间结构的来源,且无需额外的标注成本。因此,我们提出了RAIDAL(冗余感知信息密度主动学习),该方法重新利用CTC解码器,将基于表示的评分限制在解码器对齐的词汇区域,而不是让采集函数暴露于整个未过滤的视频。在三个数据集和两种架构上,RAIDAL在大词汇量、预算有限的设置中相对于竞争基线取得了最强的数据效率提升,同时在较小词汇量、大预算的设置中保持竞争力。本工作使用的代码可在此HTTP URL公开获取。

英文摘要:

Continuous sign language recognition (CSLR) is a key technology for accessibility, yet its development remains limited by the high cost of annotating continuous video streams. Active learning offers a path toward mitigating this cost, but standard acquisition functions are not designed for weakly aligned sign language videos, where sign executions are interleaved with rest poses, irregular pauses, sign-like motion, and temporally redundant frames. This temporal redundancy can undermine sample selection, as acquisition scores may be influenced by timesteps from regions that are not associated with the decoded gloss sequence, distorting the video's estimated informativeness. In this work, we show that modern CSLR models already contain a mechanism for identifying gloss-level temporal evidence: the CTC decoder. Although typically used only during inference, its alignment peaks indicate where the model localizes each predicted gloss in the feature sequence, providing a source of temporal structure for active learning acquisition functions at zero additional labeling cost. Thus, we introduce RAIDAL (Redundancy-Aware Information Density Active Learning), which repurposes the CTC decoder to restrict representation-based scoring to decoder-aligned gloss regions, rather than exposing the acquisition function to the entire unfiltered video. Across three datasets and two architectures, RAIDAL achieves its strongest data-efficiency gains over competing baselines in large-vocabulary, budget-limited settings, while remaining competitive in the smaller-vocabulary, large-budget setting. The code used in this work is publicly available at github.com/verlab/RAIDAL.

补充信息

↑