发表机构
Mila - Quebec AI Institute; McGill University; Google DeepMind; Canada CIFAR AI Chair(米拉-魁北克人工智能研究所; 麦吉尔大学; 谷歌DeepMind; 加拿大CIFAR人工智能主席项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对长文档摘要与多语言语音翻译的研究缺口,提出联合语音摘要与翻译任务,构建首个多语言跨语言基准VoxSumm,并评估相关模型表现,为多语言语音处理系统提供支撑。
AI 中文摘要
随着信息跨越语言边界的需求日益增长,用户需要对长篇内容生成简洁的跨语言表示。然而,长文档摘要研究仍以文本为中心,而多语言语音研究大多优先关注翻译,侧重保留源内容而非压缩内容。本研究通过形式化联合语音摘要与翻译(JSumT)任务来解决这一方法学缺口:即直接从源语言的长篇口语文档生成简洁、忠实的目标语言摘要。此外,本研究引入了VoxSumm,这是该任务的首个多语言及跨语言基准,包含24种语言的10045组BBC文章-摘要对,涵盖约703小时的语音数据。对代表性语音-语言模型的评估显示,不同模型及生成设置存在显著差异:Gemini3.1-Pro表现出最高的一致性,摘要生成至英语的效果总体优于生成至非英语目标语言,且先翻译整个文档再进行摘要的做法会加剧指令遵循失败。通过发布VoxSumm,本研究为开发和评估能够联合解读、压缩及翻译长篇语音的多语言系统奠定了基础。
英文摘要
As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.