AI 中文总结
RoBoSTAR提出一种结合逐部分有限标量量化与下一尺度自回归的手语生成框架,通过从粗到细的时间分辨率生成动作并并行预测身体和手部令牌,实现面向人形机器人的可重定向手语翻译。
AI 中文摘要
公共交流中的手语翻译依赖合格的专业译员,且难以规模化,这促使机器人手语成为一种补充性的无障碍交互界面。我们提出RoBoSTAR,一个以文本为条件的手语生成(SLP)框架,用于生成以人为中心的手语动作,该动作可被重定向至机器人执行,并可通过外部自动语音识别(ASR)前端可选地支持语音。传统的自回归方法将动作展平为单一全分辨率令牌序列,迫使长距离和局部依赖在统一的时间粒度上建模。RoBoSTAR则结合了逐部分有限标量量化与下一尺度自回归,在逐步更细的时间分辨率上生成动作,同时在每一步内并行预测同步的身体和手部令牌。这种从粗到细的公式在逐步细化动作细节之前提供了紧凑的长距离上下文,而自条件化和上下文损坏则增强了对跨尺度预测错误的鲁棒性。生成的动作随后被重定向至物理人形执行。我们进行了广泛的定性和定量评估,以展示RoBoSTAR的有效性。
英文摘要
Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign motion that can be retargeted for robotic execution, with speech supported optionally through an external ASR front end. Conventional autoregressive approaches flatten motion into a single full-resolution token sequence, forcing long-range and local dependencies to be modeled at a uniform temporal granularity. RoBoSTAR instead combines part-wise Finite Scalar Quantization with next-scale autoregression, generating motion over progressively finer temporal resolutions while predicting synchronized body and hand tokens in parallel within each step. This coarse-to-fine formulation provides compact long-range context before progressively refining motion details, while self-conditioning and context corruption improve robustness to cross-scale prediction errors. The generated motion is subsequently retargeted for physical humanoid execution. Extensive qualitative and quantitative evaluations are conducted to demonstrate the effectiveness of RoBoSTAR.