打破巴别塔:用于长篇字幕翻译的自进化多智能体系统
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
浏览论文内容
中文总结 AI 辅助
提出自进化多智能体系统SMART,通过动态路由、多智能体混合及裁判优化循环,实现长篇字幕翻译,并在Subtitle Arena和MuSC基准上取得最优性能。
中文摘要 AI 辅助
长篇字幕翻译需要在跨剧集或整个系列的话语和文化语境中进行推理,同时保持术语和风格的一致性。现有的单一大语言模型方法大多停留在句子级别,而多智能体系统往往采用静态工作流,无法适应场景复杂性或制作语境。我们提出SMART,一个用于长篇字幕翻译的自进化多智能体系统。在测试时训练阶段,SMART构建持久的系列级记忆,并通过动态路由器和多智能体混合层翻译部分句子,该层配备术语验证、字幕约束验证和上下文检索工具。一个裁判-优化器循环对候选翻译进行评分,并使用文本批评来更新智能体提示和路由策略,而无需重新训练底层大语言模型。在测试时推理阶段,进化后的配置翻译剩余的系列内容。我们还引入了Subtitle Arena,涵盖14种类型、每系列2至198集、制作年份1959至2023年以及15个目标语言区域,以及SubMQM,一个适应字幕的MQM框架,包含七个维度和19个错误类别。SMART在所有15个Subtitle Arena方向中取得了最佳的整体MQM分数,将平均惩罚比最强的竞争智能体系统降低了6.9%。在公开的MuSC基准上,SMART在所有四种语言对上取得了最佳模型结果,并取得了最佳人工评估结果,总体得分为4.50/5。
英文摘要
Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On a public benchmark, MuSC, SMART obtains the best model result across all 4 language pairs. SMART also achieves the best result in human evaluation with an overall score of 4.50/5.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。