arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoxSumm:用于联合摘要与翻译的多语言长篇口语新闻语料库

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Yejin Jeon, Marie Maltais, Virginia Ceccatelli, Min Ma, David Ifeoluwa Adelani

arXiv 2608.10359首次发表:更新:

发表机构

Mila - Quebec AI Institute; McGill University; Google DeepMind; Canada CIFAR AI Chair(米拉-魁北克人工智能研究所; 麦吉尔大学; 谷歌DeepMind; 加拿大CIFAR人工智能主席项目)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对长文档摘要与多语言语音翻译的研究缺口,提出联合语音摘要与翻译任务,构建首个多语言跨语言基准VoxSumm,并评估相关模型表现,为多语言语音处理系统提供支撑。

AI 中文摘要

随着信息跨越语言边界的需求日益增长,用户需要对长篇内容生成简洁的跨语言表示。然而,长文档摘要研究仍以文本为中心,而多语言语音研究大多优先关注翻译,侧重保留源内容而非压缩内容。本研究通过形式化联合语音摘要与翻译(JSumT)任务来解决这一方法学缺口:即直接从源语言的长篇口语文档生成简洁、忠实的目标语言摘要。此外,本研究引入了VoxSumm,这是该任务的首个多语言及跨语言基准,包含24种语言的10045组BBC文章-摘要对,涵盖约703小时的语音数据。对代表性语音-语言模型的评估显示,不同模型及生成设置存在显著差异:Gemini3.1-Pro表现出最高的一致性,摘要生成至英语的效果总体优于生成至非英语目标语言,且先翻译整个文档再进行摘要的做法会加剧指令遵循失败。通过发布VoxSumm,本研究为开发和评估能够联合解读、压缩及翻译长篇语音的多语言系统奠定了基础。

英文摘要

As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑