发表机构
ASELSAN; Hacettepe University(阿塞尔桑(土耳其国防工业公司); 哈杰泰佩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究将受常微分方程启发的高阶数值积分格式应用于迭代手语翻译,提出参数高效的解码器,在不增参数下提升性能,在两个基准上优于IPSLT基线。
AI 中文摘要
手语翻译借助Transformer架构已取得优异成果,但近期的改进主要依赖于扩大模型容量,代价是计算量增加。我们提出一种参数高效的替代方案,在不增加模型规模的情况下提升表达能力。我们不关注容量扩展,而是聚焦于增强迭代精化解码器的更新动态,其中每个精化步骤对应一次内部解码器迭代,逐步改进翻译生成前的潜在表示。我们从常微分方程(ODE)视角重新解释残差精化更新,并将其替换为高阶数值积分格式,即龙格-库塔方法(RK-2和RK-4)。这些方法在每个精化步骤内执行多次函数评估,以生成更准确、更稳定的表示更新,且不增加解码器参数。据我们所知,这是受ODE启发的更新动态首次应用于手语翻译。RK-2在PHOENIX-2014-T测试集上达到22.96的BLEU-4值,在CSL-Daily测试集上达到19.34的BLEU-4值,在两个基准上均优于IPSLT基线,且在CSL-Daily上使用更少的解码器层和精化迭代。这些结果表明,更强的精化动态可在参数高效的解码器设计下提升翻译性能,为传统模型扩展提供了互补替代方案。
英文摘要
Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. We propose a parameter-efficient alternative that improves expressiveness without increasing model size. Rather than scaling capacity, we focus on enhancing the update dynamics of iterative refinement decoders, where each refinement step corresponds to one internal decoder iteration that progressively improves the latent representation before translation generation. We reinterpret residual refinement updates from an Ordinary Differential Equation (ODE) perspective and replace them with higher-order numerical integration schemes, namely Runge--Kutta methods (RK-2 and RK-4). These methods perform multiple function evaluations within each refinement step to produce more accurate and stable representation updates without adding decoder parameters. To the best of our knowledge, this is the first application of ODE-inspired update dynamics to sign language translation. RK-2 achieves 22.96 BLEU-4 on the PHOENIX-2014-T test set and 19.34 BLEU-4 on the CSL-Daily test set, outperforming the IPSLT baseline on both benchmarks, with fewer decoder layers and refinement iterations on CSL-Daily. These results suggest that stronger refinement dynamics can improve translation performance under parameter-efficient decoder designs, providing a complementary alternative to conventional model scaling.
CommentsAccepted at the 14th International Workshop on Assistive Computer Vision and Robotics (ACVR 2026), held in conjunction with ECCV 2026