AI 中文总结
针对现有翻译模型难以处理日常对话现象的问题,本文提出多语种对话语料库BaatCheet,含约49,000个对话,覆盖五个翻译方向,并微调五个开源LLM,证明微调显著优于零样本和少样本基线,采用多种评估策略验证翻译质量。
AI 中文摘要
现有的翻译模型通常基于句子级和正式文本进行训练,这限制了它们捕捉日常对话现象(如非正式性、说话者互动和话语连贯性)的能力。大多数现有的印度语言翻译资源和评估基准都侧重于句子级或正式文本,使得评估对话现象的翻译质量变得困难。在这项工作中,我们引入了BaatCheet,一个以印地语中表示对话或闲聊的术语命名的多语种对话语料库,包含约49,000个对话,覆盖五个翻译方向。我们微调了五个开源大型语言模型,并使用了七种训练数据配置,发现微调相比零样本和少样本基线带来了显著的性能提升。为了全面评估对话翻译质量,我们采用了多种评估策略,包括自动指标、LLM作为评判者,以及使用基于SQM引导的直接评估(DA)协议进行的人工评估。
英文摘要
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level or formal text, making it difficult to assess translation quality of the dialogue phenomena. In this work, we introduce BaatCheet, a multilingual dialogue corpus named after the Hindi term for conversation or chitchat, containing approximately 49,000 dialogues for dialogue translation across five translation directions. We fine-tune five open-source LLMs across seven training data configurations and find that fine-tuning yields substantial gains over zero- and few-shot baselines. To comprehensively evaluate dialogue translation quality, we employ multiple evaluation strategies, including automatic metrics, LLM-as-judge, and human assessments using an SQM-guided Direct Assessment (DA) Protocol.