SignFLIP:一种通过大规模阶段式对齐实现手语翻译与生成的统一模型
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
浏览论文内容
中文总结 AI 辅助
SignFLIP提出以LLM为中心的对称架构和阶段式训练策略,在大规模数据上实现手语翻译与生成的统一双向对齐,并在多个基准上达到与任务特定模型相当的性能且强迁移至手语识别。
中文摘要 AI 辅助
手语翻译和生成共享文本与手语表示之间双向对齐的目标。然而,现有方法要么将它们视为孤立的任务,要么仅在有限的数据集上得到验证,限制了模态间有效建模。在本文中,我们提出SignFLIP,一个以LLM为中心的统一翻译与生成框架。为了实现文本与手语之间的双向映射,SignFLIP采用对称架构以及基于大规模数据的阶段式训练策略。共享的手语-文本表示被逐步细化:预对齐促进后续的手语翻译(SLT),而经SLT适配的表示进一步有利于手语生成(SLG)。在多个基准上的大量实验表明,SignFLIP在翻译和生成任务上均展现出与任务特定模型相当的性能,并且对手语识别具有强迁移性。
英文摘要
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.
发表机构
- Institute of Science Tokyo(东京科学大学)
- Chungbuk National University(忠北国立大学)
机构由 AI 辅助整理,请以论文原文为准。