arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11131cs.CL

对齐评分量表的同声传译解耦评估

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

  • School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)
  • School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

Ziyu Zhang, Satoshi Nakamura

AI总结:

针对人类同声传译缺乏与评分量表对齐的自动评估问题,本文构建了1,101个片段的标注语料库,并提出在LoRA适配的COMET-KIWI编码器上使用双回归头,以解耦各维度,在LQ和EXP上分别达到0.388和0.301的皮尔逊相关性。

AI中文摘要:

人类同声传译(SI)通常使用分析性评分量表进行评估,该量表将意义传递、交付质量和时间同步性分开,但目前尚无自动度量标准用于与评分量表对齐的片段级SI评估。我们构建了一个包含1,101个SI片段、带有意义传递(LQ)、交付质量(EXP)和感知延迟(LAT)评分的专业标注语料库。我们表明,结构化LLM提示和标量监督会压缩评分量表维度,导致与人类评分的相关性接近零,并且跨维度耦合严重。为了在相同骨干容量下隔离监督结构,我们在LoRA适配的COMET-KIWI编码器上引入了双回归头。在留出的谈话级测试集上,该模型在LQ上达到0.388、在EXP上达到0.301的皮尔逊相关系数,优于冻结的COMET-KIWI。鉴于绝对评分者一致性较低,我们将结果相对于人类一致性进行解释,并针对形成性评估的目标稳定排序信号。

英文摘要:

Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annotated corpus of 1,101 SI segments with scores for meaning transfer (LQ), delivery quality (EXP), and perceived latency (LAT). We show that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling. To isolate supervision structure under identical backbone capacity, we introduce dual regression heads on a LoRA-adapted COMET-KIWI encoder. On a held-out talk-level test set, the model achieves Pearson correlations of 0.388 (LQ) and 0.301 (EXP), improving over frozen COMET-KIWI. Given low absolute rater agreement, we interpret results relative to human consistency and target stable ranking signals for formative assessment.

↑