双流同步翻译:基于二维网格注意力
Dual-Stream Simultaneous Translation via 2D Grid Attention
浏览论文内容
中文总结 AI 辅助
提出双流网格注意力框架,通过四种注意力类型联合归一化建模源目标流交互,以两种近似降低复杂度,在汉英同步翻译中显著提升质量并降低延迟。
中文摘要 AI 辅助
同步机器翻译必须在源输入完整之前生成目标词元。现有方法通过事后读写策略解决这一问题,使注意力机制无法感知双向流依赖。我们提出一种双流注意力框架,将源流和目标流表示为二维隐藏状态网格,并通过联合QK Softmax归一化合并的四种结构不同的注意力类型来建模它们的交互。两种近似方法——广播和哈达玛积——将每层复杂度从O(X^2Y+XY^2)降低到O(X^2+Y^2+XY),且误差可证明衰减。训练采用自引导循环:逐单元损失热图驱动动态规划路径恢复,生成读写决策监督标签,无需外部对齐。增量KV缓存配合锚定旋转位置嵌入实现高效流式推理。在汉英同步翻译中,所提模型在相似延迟下比Wait-k基线高出+5.66 BLEURT和+10.36 COMET,并在COMET上以响应延迟的一小部分超越非流式参考。
英文摘要
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
发表机构
- Tsinghua University(清华大学)
- Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学具身智能与机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。