arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习标签噪声下的手语识别:隔离与连续设置下噪声鲁棒损失函数的研究

Learning Sign Language Recognition under Label Noise: A Study of Noise-Robust Losses for Isolated and Continuous Settings

Akihisa Shitara, Yoichi Ochiai

arXiv 2609.12885首次发表:更新:

发表机构

University of Tsukuba; Pixie Dust Technologies, Inc.(筑波大学; Pixie Dust Technologies 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对手语识别中标签噪声问题,将图像分类的鲁棒损失(SCE、GCE)应用于单标签分类器,在隔离与连续设置下评估其性能,发现SCE设置通用但波动大,连续设置增益受辅助权重影响。

AI 中文摘要

在手语识别中,隔离手语识别(ISLR)的分类损失将单个标签视为真实标签,我们为连续手语识别(CSLR)方法添加的基于伪标签的帧级辅助分类器也是如此,而CSLR方法本身缺乏此类标签。风格差异使ISLR标注变得模糊,而CSLR缺乏时间边界则迫使使用伪标签;两者都存在噪声。因此,我们将图像分类中交叉熵(CE)的鲁棒替代方案——对称交叉熵(SCE)和广义交叉熵(GCE)——应用于前面的单标签分类器,而非连接主义时间分类。在ASL Citizen数据集上,使用三个骨干网络(ST-GCN使用三个种子)注入对称噪声,当标签干净时,鲁棒损失最多损失2.5个百分点,在噪声率为0.2的所有六种条件下,比CE高出2.9-10.0个百分点,其中一种条件仅在开发集上重新选择q后才达到此效果。GCE提升更大,但其最优q不能跨骨干网络迁移,而一种SCE设置在所有九种条件下均有效;两者在多次运行中的波动比CE大2-11倍,因此有利的点估计并不能确立稳定性。对于CSLR(PHOENIX-2014),我们报告没有增益;我们的帧级目标带有系统性分配偏差,使该研究成为对单一配置的诊断。在lambda_aux=25时,伪标签CE辅助在VAC、CorrNet和SlowFastSign上将词错误率提高到无辅助基线之上,而GCE/SCE比CE改进1.7-3.2个百分点(六种条件中有三种回到该基线以下)。然而,在共同的lambda_aux下,三种损失的有效梯度差异超过一个数量级:匹配初始梯度将差距缩小到0.4-0.9个百分点,仅降低CE权重就已超过该基线,因此退化与改进都无法与权重的影响分离。我们仅使用对称噪声;多种子评估仅覆盖ST-GCN和VAC隔离设置。

英文摘要

In sign language recognition, the isolated (ISLR) classification loss treats a single label as ground truth, as does the frame-level auxiliary classifier over pseudo-labels we add to continuous (CSLR) methods, which lack one. Stylistic variation blurs ISLR annotation and the lack of temporal boundaries in CSLR forces pseudo-labels; both are noisy. We therefore apply symmetric and generalized cross entropy (SCE, GCE), robust alternatives to cross entropy (CE) from image classification, not to connectionist temporal classification but to the preceding single-label classifier. On ASL Citizen with injected symmetric noise on three backbones (three seeds for ST-GCN), robust losses cost at most 2.5 pt when labels are clean and beat CE by 2.9-10.0 pt in all six conditions at noise rate 0.2, one of which only after q was re-selected on dev. GCE gains more, but its optimal q does not transfer across backbones, whereas one SCE setting works in all nine conditions; both vary 2-11 times more than CE across runs, so a favorable point estimate does not establish stability. For CSLR (PHOENIX-2014) we report no gain; our frame-level targets carry a systematic assignment bias, making that study a diagnosis of a single configuration. At lambda_aux = 25 the pseudo-label CE auxiliary raises word error rate above the no-auxiliary baseline on VAC, CorrNet and SlowFastSign, and GCE/SCE improve on CE by 1.7-3.2 pt (three of six conditions return below that baseline). However, the three losses differ by more than an order of magnitude in effective gradient at a common lambda_aux: matching the initial gradient shrinks the gap to 0.4-0.9 pt, and lowering the CE weight alone already beats that baseline, so neither the degradation nor the improvement can be separated from the effect of the weight. We use only symmetric noise; multi-seed evaluation covers only ST-GCN and VAC isolated.

Comments15 pages, 1 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑