arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25493cs.CV

SMART:用于统一手语识别与定位的多模态大语言模型(MLLM)引导的时间对齐框架

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

  • Dankook University(檀国大学)

机构由 AI 辅助整理,请以论文原文为准。

Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin, Younggeun Choi

AI总结:

本研究提出 MLLM 引导的 SMART 框架,整合多尺度时间适配器与 CSFormer 模块,在四个手语基准数据集上实现了手语识别与定位任务的性能提升。

AI中文摘要:

连续手语识别(CSLR)旨在在弱序列级监督下,从未分段的手语视频中识别 gloss 序列。然而,现有方法依赖句子级 gloss 标注,为细粒度表示学习提供的时间和语义指导有限。传统视频-文本对齐还需要大批次,这对于内存密集型手语视频训练而言效率低下。本研究中,我们提出 SMART,一种用于联合手语识别与定位的多模态大语言模型(MLLM)引导的时间对齐框架。SMART 利用 MLLM 生成的动作描述作为辅助语义线索,并在小批次训练下实现稳定的视频-文本对齐。为改进时间表示学习,我们引入多尺度时间适配器,在 Transformer 编码过程中建模时间交互。针对密集时间定位,SMART 整合 CSFormer,这是一种由 CSLR 引导的定位模块,可将识别得到的 gloss 证据注入边界感知定位网络。该统一框架使 CSLR 特征能够惠及定位任务,同时定位监督补充基于连接时序分类(CTC)的弱识别。在四个手语基准数据集(包括 PHOENIX14-T、CSL-Daily、大规模 KSL 以及灾害与安全 KSL 数据集)上开展的实验,证明了 SMART 在识别与定位两项任务上的有效性。

英文摘要:

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional video-text alignment also requires large batch sizes, making it inefficient for memory-intensive sign language video training. In this work, we propose SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting. SMART uses MLLMgenerated motion descriptions as auxiliary semantic cues and performs stable videotext alignment under small-batch training. To improve temporal representation learning, we introduce a Multi-Scale Temporal Adapter that models temporal interactions during transformer encoding. For dense temporal localization, SMART incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network. This unified framework enables CSLR features to benefit spotting, while spotting supervision complements weak CTC-based recognition. Experiments on four sign language benchmarks, including PHOENIX14-T, CSL-Daily, Large-scale KSL, and Disaster and Safety KSL datasets, demonstrate the effectiveness of SMART across both recognition and spotting tasks.

补充信息

↑