arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义难度并非视觉难度:用于手语检索的符号感知硬负样本挖掘

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

Junmyeong Lee, Chan Hur, ChangSu Choi, Sukmin Cho, Fitsum Gaim, Eui Jun Hwang, Hoyun Song, KyungTae Lim

arXiv 2607.09263首次发表:更新:

发表机构

School of Computing; Graduate School of Culture Technology; Korea Advanced Institute of Science and Technology; ETRI Medical Informatics Laboratory(计算机学院; 文化技术研究生院; 韩国科学技术院; 电子通信研究院医学信息学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究手语检索在细粒度场景的问题,提出符号感知硬负样本挖掘方法,通过在嵌入空间基于视觉易混淆性构建硬负样本,实验证明该方法能提升细粒度检索性能且保持粗粒度准确性。

AI 中文摘要

手语检索(SLRet)能高效访问手语内容,但在需区分视觉相似符号的细粒度场景中仍很脆弱。我们表明此限制并非源于模型能力,而是无效的硬负样本监督。具体而言,我们将细粒度检索失败表述为负分布不匹配:语义不同但视觉上易混淆的符号很少被视为硬负样本,而现有基于文本的挖掘策略无法捕捉这种视觉模糊性。为解决此问题,我们提出符号感知硬负样本挖掘(SAN),它基于手语嵌入空间中的视觉易混淆性构建硬负样本。在PHOENIX - 2014T上的实验表明,SAN在保持粗粒度准确性的同时显著提高了细粒度检索性能,突出了在手语检索中使负样本监督与视觉模糊性对齐的重要性。

英文摘要

Sign Language Retrieval (SLRet) enables efficient access to sign language content but remains fragile in fine-grained scenarios where visually similar signs must be distinguished. We show that this limitation does not stem from model capacity, but from ineffective hard negative supervision. Specifically, we formulate fine-grained retrieval failures as a negative distribution mismatch: semantically distinct yet visually confusable signs are rarely treated as hard negatives, while existing text-based mining strategies fail to capture such visual ambiguity. To address this issue, we propose Sign-Aware Hard Negative Mining (SAN), which constructs hard negatives based on visual confusability in the sign embedding space rather than linguistic similarity. Experiments on PHOENIX-2014T demonstrate that SAN substantially improves fine-grained retrieval performance while preserving coarse-grained accuracy, highlighting the importance of aligning negative supervision with visual ambiguity in sign language retrieval.

CommentsAccepted to ACL 2026 main

Journal refProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pages 28262-28277

DOI:10.18653/v1/2026.acl-long.1302

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑