请再翻译一次:关于类人推理是否有助于翻译的两个简单实验
Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation
- Language Technology Lab University of Amsterdam(语言技术实验室 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过实验检验LLM翻译中类人逐步推理的有效性,发现CoT式分解并非性能提升的明确来源,而“再翻译一次”的自我精炼策略效果更优,揭示了人类与LLM最优翻译策略的差异。
AI中文摘要:
大型语言模型(LLMs)在许多任务中展现出强大的推理能力,通常通过思维链(Chain-of-Thought, CoT)推理显式地分解任务。近期基于LLM的翻译工作设计了手工构造的提示来分解翻译过程,或训练模型引入中间步骤。例如,逐步翻译(Translating Step-by-step,Briakou et al., 2024)引入了一种多步提示,利用LLM对翻译进行分解与精炼,并在WMT24测试数据上取得了最先进的结果。在本工作中,我们审视了这一策略的有效性。从实证结果看,我们没有发现明确证据表明性能提升源于通过CoT显式分解翻译过程,至少对于所测试的模型而言如此;并且我们表明,提示LLM“再翻译一次”并进行自我精炼,其结果甚至优于类人的逐步提示。虽然分解会影响翻译行为,但对分解的忠实执行对翻译既有正面影响也有负面影响。因此,我们的分析表明,适用于人类和LLM的最优翻译策略之间存在分歧。
英文摘要:
Large Language Models (LLMs) demonstrate strong reasoning capabilities for many tasks, often by explicitly decomposing the task via Chain-of-Thought (CoT) reasoning. Recent work on LLM-based translation designs hand-crafted prompts to decompose translation, or trains models to incorporate intermediate steps. Translating Step-by-step (Briakou et al., 2024), for instance, introduces a multi-step prompt with decomposition and refinement of translation with LLMs, which achieved state-of-the-art results on WMT24 test data. In this work, we scrutinise this strategy's effectiveness. Empirically, we find no clear evidence that performance gains stem from explicitly decomposing the translation process via CoT, at least for the models on test; and we show prompting LLMs to 'translate again' and self-refine yields even better results than human-like step-by-step prompting. While the decomposition influences translation behaviour, faithfulness to the decomposition has both positive and negative effects on translation. Our analysis therefore suggests a divergence between the optimal translation strategies for humans and LLMs.