arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

APEX-VW:医疗领域的文档级英西后编辑数据集

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Marie Escribe, Tharindu Ranasinghe, Amal Haddad Haddad, Hansi Hettiarachchi, Damith Premasiri

arXiv 2608.08059首次发表:更新:

发表机构

Universitat Politècnica de València; Lancaster University; Universidad de Granada(瓦伦西亚理工大学; 兰卡斯特大学; 格拉纳达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出APEX-VW医疗领域文档级英西后编辑数据集,含4.2万词文档及专业后编辑数据,可用于术语规范化等研究,为相关任务提供基准。

AI 中文摘要

机器翻译(MT)输出的后编辑(PE)常需在多个片段中重复相同词汇和术语修正,尤其在专业且高度重复的文档中。尽管自动后编辑(APE)领域已有大量研究,但多数可用语料库为句子级,其余为合成数据,且整体并非用于研究现实计算机辅助翻译(CAT)工作流中修正的传播规律。本文提出APEX-VW(Virtual Wards自动后编辑实验)语料库,这是一个全新的开放英西(EN-ES)数据集,由近期NHS虚拟病房文档及Trados Studio中的专业后编辑工作构建,具备受控的机器翻译、术语和质量保证设置。该语料库包含7个文档连贯的源文本,总计4.2万词,经4个代表不同范式的机器翻译系统翻译后,由专业译员进行后编辑。与WMT APE语料库、eSCAPE、MLQE-PE或LangMark等现有资源不同,该数据集保留了文档顺序和CAT工具上下文,适用于术语规范化、修正传播及人在回路翻译支持的研究。本文描述了语料库设计、数据准备、后编辑设置及初始语料库统计,并将该资源定位为文档级APE和感知传播辅助工具的基准。

英文摘要

Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑