arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29131cs.CLcs.AI

标签感知的结构化文本翻译:迈向系统性理解

Tag-Aware Structured Text Translation: Towards a Systematic Understanding

Zhanglin Wu, Hengchao Shang, Daimeng Wei, Jiaxin Guo, Zongyao Li, Tengfei Song, Ning Xie, Weidong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对带标签文本翻译中流畅性与标签保真度的矛盾,提出数据合成、能力构建、多目标对齐三层面系统方法,显著提升多语言翻译质量。

中文摘要 AI 辅助

互联网文本中充斥着承载结构、语义和功能意义的格式标签。当前基于大语言模型(LLM)的翻译系统在处理带标签文本时,难以在翻译流畅性与标签保真度之间取得平衡。我们认为,解决这一矛盾需要在三个相互关联的层面采取系统性方法:数据合成、能力构建和多目标对齐。在数据层面,我们识别并形式化了合成数据生成中结构标签多样性与翻译自然性之间的根本性权衡;现有方法往往以牺牲其中一方为代价来优化另一方。我们提出了一种混合合成策略(Hy-LST),结合基于LLM的合成标签方法和基于LLM的两阶段合成标签方法,以生成既多样又自然的带标签数据。在能力层面,我们将标签感知翻译分解为多任务监督微调框架中难度递增的四个子任务,从而实现针对性的能力获取和知识迁移。在对齐层面,我们在组相对策略优化框架下设计了三个互补的奖励函数,每个函数针对不同的目标(流畅性、标签保真度和标签范围内的翻译质量),并证明联合优化始终优于单一奖励的替代方案。在六个语言方向(英中、英日、英德、英法、英俄、德法)上的实验表明,每个层面都带来了可衡量的改进,完整系统显著优于现有方法。定性分析揭示了使用我们方法训练后出现的特定错误模式及其缓解情况。

英文摘要

Internet texts are replete with format tags that carry structural, semantic, and functional meaning. Current large language model (LLM)-based translation systems struggle to balance translation fluency with tag fidelity when processing tagged text. We argue that resolving this tension requires a systematic approach at three interconnected levels: data synthesis, capability building, and multi-objective alignment. At the data level, we identify and formalize a fundamental trade-off between structural tag diversity and translation naturalness in synthetic data generation; existing methods optimize for one at the expense of the other. We propose a hybrid synthesis strategy (Hy-LST) combining LLM-based synthesis tag method and Two-Stage LLM-based synthesis tag method to produce both diverse and natural tagged data. At the capability level, we decompose tag-aware translation into four sub-tasks of increasing difficulty in a multi-task supervised fine-tuning framework, enabling targeted capability acquisition and knowledge transfer. At the alignment level, we design three complementary reward functions under a group relative policy optimization framework, each targeting a distinct objective (fluency, tag fidelity, and tag-scoped translation quality), and show that joint optimization consistently outperforms single-reward alternatives. Experiments on six language directions (en2zh, en2ja, en2de, en2fr, en2ru, de2fr) demonstrate that each level contributes measurable improvements, and the complete system significantly outperforms existing methods. Qualitative analysis reveals specific error patterns and their mitigation after training with our method.

发表机构

  • Huawei Translation Service Center(华为翻译服务中心)

机构由 AI 辅助整理,请以论文原文为准。

↑