发表机构
University of Amsterdam; Vrije Universiteit Amsterdam(阿姆斯特丹大学; 阿姆斯特丹自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究LLM的序列式与并行式测试时缩放在机器翻译中的特性,发现序列式采样性能上限更高,可提升翻译流畅度但在大推理预算下会降准,还分析了其机制与鲁棒性。
AI 中文摘要
大语言模型(LLM)的两种测试时缩放形式已成为有效且被广泛采用的范式:序列式,即后续的答案尝试依赖于先前的尝试;以及并行式,例如带重排序的独立同分布(i.i.d.)采样。本研究调查了它们在翻译任务中的特性。首先,我们的研究表明,序列式采样具有更高的性能上限,能提供更多样、更有效的样本池,尤其在较小的采样预算下表现突出。其次,我们通过多维度人工分析探究测试时缩放的本质,对N-best翻译的人工分析显示,序列式采样可显著提升翻译的流畅性与自然度,但在推理预算较大时会降低翻译准确性。最后,我们提出了序列式缩放提升机器翻译性能的机制解释,受控分析将序列式自改进的成功部分归因于模型能访问更大的目标侧上下文。对序列式采样的消融实验表明,其在不同采样温度下具有鲁棒性,同时也揭示了其对上下文构建的敏感性,为未来改进指明了方向。
英文摘要
Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective and widely adopted paradigms: sequential, in which later answer attempts depend on earlier ones, and parallel, such as i.i.d. sampling with reranking. In this study, we investigate their properties in translation. First, our study shows that sequential sampling has a higher performance ceiling, providing a more diverse and effective pool of samples, particularly under smaller sampling budgets. Second, we interrogate the nature of test-time scaling through a multidimensional manual analysis. Human analysis of the Best-of-N translations demonstrates that sequential sampling substantially improves translation fluency and naturalness, but can degrade accuracy when inference budgets are large. Finally, we suggest an explanation of the mechanism through which sequential scaling improves machine translation. Our controlled analysis partially attributes the success of sequential self-improvement to the model's access to a larger target-side context. Ablation experiments on sequential sampling demonstrate its robustness across different sampling temperatures, while also revealing sensitivity to context construction, suggesting directions for future improvement.
CommentsAccepted to Findings of EMNLP 2026