发表机构
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对连续手语视频分割中句子边界模糊的问题,提出差异感知框架SignShift,通过时间差异模块建模帧间特征变化并预测句子数量,显著提升分割性能。
AI 中文摘要
手语理解的最新进展在短时、单句视频上取得了令人瞩目的成功,然而当应用于长时、连续的手语视频时,其性能急剧下降。为弥合这一差距,我们聚焦于一个具有挑战性且现实的环境:纯视觉句子级手语分割(Vis-SSLS),其目标是在没有任何字幕辅助的情况下,将连续的手语视频划分为不重叠的句子级片段,作为下游识别和翻译任务的关键前提。然而,手语中的句子转换通常平滑且在视觉上模糊,缺乏明确的停顿或姿态重置。因此,静态帧表示可能无法捕捉指示句子边界的细微时间变化。为解决这一挑战,我们提出了SignShift,一个差异感知分割框架,明确地将帧间特征变化建模为句子边界检测的语义线索。首先,为建模特征变化,我们设计了一个时间差异模块,该模块整合了全帧、面部和手部线索,并采用帧间差分来学习多尺度时间变化,从而捕捉细粒度的局部运动学和全局语义转换。其次,为缓解过分割和欠分割问题,我们设计了一个片段计数预测模块,该模块预测句子数量以指导边界选择。在基准数据集上的大量实验表明,SignShift显著优于现有方法,验证了其有效性。
英文摘要
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.