arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

看见语义转变:手语视频的差异感知句子级时间分割

Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos

Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Kuizhuang Liu, Zhiwei Jiang, Lei Xie

arXiv 2609.31148首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对连续手语视频分割中句子边界模糊的问题,提出差异感知框架SignShift,通过时间差异模块建模帧间特征变化并预测句子数量,显著提升分割性能。

AI 中文摘要

手语理解的最新进展在短时、单句视频上取得了令人瞩目的成功,然而当应用于长时、连续的手语视频时,其性能急剧下降。为弥合这一差距,我们聚焦于一个具有挑战性且现实的环境:纯视觉句子级手语分割(Vis-SSLS),其目标是在没有任何字幕辅助的情况下,将连续的手语视频划分为不重叠的句子级片段,作为下游识别和翻译任务的关键前提。然而,手语中的句子转换通常平滑且在视觉上模糊,缺乏明确的停顿或姿态重置。因此,静态帧表示可能无法捕捉指示句子边界的细微时间变化。为解决这一挑战,我们提出了SignShift,一个差异感知分割框架,明确地将帧间特征变化建模为句子边界检测的语义线索。首先,为建模特征变化,我们设计了一个时间差异模块,该模块整合了全帧、面部和手部线索,并采用帧间差分来学习多尺度时间变化,从而捕捉细粒度的局部运动学和全局语义转换。其次,为缓解过分割和欠分割问题,我们设计了一个片段计数预测模块,该模块预测句子数量以指导边界选择。在基准数据集上的大量实验表明,SignShift显著优于现有方法,验证了其有效性。

英文摘要

Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑