OneSign:用一个模型统一手语理解任务
OneSign: Unifying Sign Language Understanding Tasks with One Model
浏览论文内容
中文总结 AI 辅助
OneSign提出统一框架,通过模态自适应混合专家架构,在单一模型中同时处理孤立词识别、连续识别和翻译任务,实现多任务统一并达到先进性能。
中文摘要 AI 辅助
手语理解(SLU)涵盖多种任务,包括孤立词手语识别(ISLR)、连续手语识别(CSLR)和手语翻译(SLT)。尽管这些任务共享基本的语义和语言基础,但它们通常采用特定于任务的架构和训练流程来处理,这阻碍了知识共享,并需要为每个任务进行昂贵的预训练和微调。在本文中,我们关注SLU任务的两个方面:(1)训练和推理流程高度碎片化:大多数方法依赖于在大规模SL数据集上进行预训练,然后进行任务或数据集特定的微调,这导致产生多个专门模型而非单个检查点。(2)当前基于LLM的方法可能忽略手语和文本标记之间固有的模态差异,简单地将它们拼接起来,并用相同的解码器层处理两种模态。在本文中,我们提出了OneSign,一个统一框架,在单个模型和单个检查点内处理多个SLU任务。OneSign在单一训练范式下重新定义了ISLR、CSLR和SLT。为了适应手语和文本表示的异质特征,我们引入了一种模态自适应混合专家(MA-MoE)架构,由共享专家和针对手语和文本标记的模态特定专家组成。模态路由器动态激活相应的专家,其输出被聚合以形成最终的标记表示。通过实现依赖模态的专家专业化,同时保留共享专家路径,MA-MoE可以有效地建模连续手语表示和离散文本标记之间的模态差异。在多个基准上的大量实验表明,OneSign在几个基准上取得了具有竞争力或最先进的性能,突显了其作为统一SLU模型的有效性。数据集可在以下网址获取:此https URL。
英文摘要
SLU encompasses a diverse set of tasks, including ISLR, CSLR, and SLT. Although these tasks share basic semantic and linguistic foundations, they are typically addressed with task-specific architectures and training pipelines, which hinders knowledge sharing and requires costly pretraining and finetuning for each task. In this paper, we focus on two aspects of SLU tasks: (1) training and inference pipelines are highly fragmented: most methods rely on pretraining on large-scale SL datasets followed by task- or dataset-specific finetuning, which leads to multiple specialized models rather than a single checkpoint. (2) current LLM-based methods may overlook the inherent modality discrepancy between sign and text tokens, simply concatenating them and processing both modalities with the same decoder layers. In this paper, we present OneSign, a unified framework that addresses multiple SLU tasks within a single model and a single checkpoint. OneSign reformulates ISLR, CSLR, and SLT under a single training paradigm. To accommodate the heterogeneous characteristics of sign and text representations, we introduce a Modality-Adaptive Mixture-of-Experts (MA-MoE) architecture, consisting of a shared expert and modality-specific experts for sign and text tokens. A modality router dynamically activates the corresponding experts, and their outputs are aggregated to form the final token representations. By enabling modality-dependent expert specialization while preserving a shared expert path, MA-MoE can effectively model the modality differences between continuous sign representations and discrete text tokens. Extensive experiments on multiple benchmarks demonstrate that OneSign achieves competitive or state-of-the-art performance on several benchmarks, highlighting its effectiveness as a unified SLU model. Datasets are available at : https://github.com/gswycf/OneSign.