发表机构
ShanghaiTech University; DGene; University of Derby(上海科技大学; DGene; 德比大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReSCUE提出统一框架,通过推理感知训练、稳定重翻译和句子承诺机制,实现未分段长形式手语视频的同步翻译,在低延迟下接近离线系统质量。
AI 中文摘要
同步手语翻译(SLT)对于实时通信至关重要,然而现有方法大多局限于句子级别、离线设置,并假设输入已预先分段。这些假设阻碍了在涉及连续、未分段视频流的现实场景中的部署。我们提出了ReSCUE,一个用于未分段长形式手语视频的同步SLT统一框架,该框架将训练和推理与现实的流式条件对齐。ReSCUE结合了推理感知训练以处理部分输入、非手语停顿和多句子上下文,稳定的重翻译以实现低延迟但可修正的预测并减少输出闪烁,以及一个句子承诺机制用于在线分段和内存管理。在标准句子级基准上的实验表明,ReSCUE在低延迟设置下实现了更低的延迟和最佳的翻译质量。在长形式未分段数据集上,ReSCUE接近使用真实句子边界的oracle离线系统的翻译质量,同时以显著更低的延迟运行,展示了其在现实世界流式场景中的实用性。
英文摘要
Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.
CommentsAccepted at NeurIPS 2026