发表机构
Huawei Translation Service Center; School of Informatics, Xiamen University(华为翻译服务中心; 厦门大学信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EviSI,一种基于大语言模型的同声传译评估智能体,采用MQM原则评估语义与口语表达,在英译中方向上恢复人工系统排名,平均Kendall一致性达0.707。
AI 中文摘要
同声语音到语音翻译要求在源语音流持续进行的同时完成理解、翻译和口语表达。为支持及时交付并限制累积延迟,系统采用重构和摘要化策略,这些策略在偏离书面参考的同时能够保留语义。BLEU和COMET可能无法可靠地将这种变异与语义损失区分开来。我们提出EviSI,一个基于大语言模型的评估智能体,它借鉴了多维质量指标(MQM)的错误分析和扣分原则。该智能体构建共享的源证据,评估语义忠实度和口语表达,协调重叠错误并以确定性方式评分。EviSI恢复了英译中方向上的总体人工系统排名。在语料库内,与人工系统排名的平均Kendall一致性在英译中方向上达到0.707,在中译英方向上达到0.467,超过了所评估的基线。一项涵盖五个方向的扩展实验显示,在没有人工评分的情况下,与COMET呈正相关。与人工的个体输出一致性结果仍参差不齐。
英文摘要
Low-latency simultaneous speech-to-speech translation must keep pace with ongoing speech while preserving key information. To meet these demands, systems use segmentation, reformulation and condensation to reorganize and rephrase information. However, metrics developed for text translation, including BLEU and COMET, may not consistently distinguish faithful adaptations from semantic errors. We propose EviSI, a large language model evaluation agent combining Multidimensional Quality Metrics (MQM) with criteria developed with professional interpreters. Shared source evidence guides assessment across four dimensions: Anchor, Event, Logic and Fluency. Verified errors are deduplicated before deterministic scoring. On human-rated English to Chinese and Chinese to English data, EviSI recovers the aggregate English to Chinese human system ranking. Mean within-dataset Kendall correlations for system rankings reach 0.707 and 0.467, respectively, exceeding evaluated BLEU and COMET baselines. A multilingual extension to five directions without human ratings retains the dimensions and scoring rule, showing positive system ranking correlations with COMET throughout.