评估机器翻译对话中的交际成功
Evaluating Communicative Success in Machine-Translated Conversation
浏览论文内容
中文总结 AI 辅助
本研究提出三层清单与评判框架,从语义、语用和文化社会维度评估机器翻译口译对话的交际成功,经多语言基准验证,发现传统保真度指标忽视交际失败,而场景与文化上下文可提升成功。
中文摘要 AI 辅助
基于机器翻译(MT)的口译智能体日益成为语言不通者之间实时对话的中介,然而我们仍使用为孤立句子设计的指标来评估它们,这些指标衡量的是保真度而非交际是否成功。我们引入了一个可复用的三层清单与评判框架,该框架从语义、语用和文化社会维度评估口译中介对话,涵盖了保真度指标未测量的自然度、意图和社会适宜性。该框架可在单轮和交互式多轮设置中运行,在交互式多轮设置中,模拟用户随着对话展开对翻译后的消息进行回复,每一轮与整个对话一起被评分。我们通过受控扰动、跨评判者比较和人工标注对其进行了广泛验证。我们的主要单轮基准评估了涵盖阿拉伯语、孟加拉语、印度尼西亚语和韩语的10种口译设置,这些设置来自跨越12个翻译方向的5,624个OpenSubtitles衍生场景,而我们的多轮研究以脚本和实时模式覆盖了全部6种语言对。结果表明,从语义到语用和文化社会成功呈持续下降趋势,而传统MT指标忽视了较强口译设置中的失败,提示消融实验表明,场景上下文、结构化指令和文化上下文能提高交际成功,尽管提升幅度因设置而异。因此,我们的工作为对话中的口译智能体提供了评估框架和基准,并强调了交际成功与现有翻译指标并重的重要性。
英文摘要
Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds. We introduce a reusable three-layer checklist-and-judge framework that evaluates interpreter-mediated conversation across semantic, pragmatic, and cultural-social dimensions, covering the naturalness, intent, and social appropriateness that fidelity metrics leave unmeasured. It runs in both single-turn and interactive multi-turn settings, where simulated users reply to translated messages as the conversation unfolds and each turn is scored alongside the conversation as a whole. We extensively validate it through controlled perturbations, cross-judge comparisons, and human annotations. Our main single-turn benchmark evaluates 10 interpreter setups across Arabic, Bengali, Indonesian, and Korean from 5,624 OpenSubtitles-derived scenarios spanning 12 translation directions, and our multi-turn study covers all 6 language pairs in scripted and live modes. Results show a consistent decline from semantic to pragmatic and cultural-social success, while conventional MT metrics overlook failures among stronger interpreters, and prompt ablations show that scenario context, structured instructions, and cultural context improve communicative success, although gains vary across setups. Our work thus provides an evaluation framework and benchmark for interpreter agents in conversation, and highlights the importance of communicative success alongside existing translation metrics.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。