发表机构
Technische Hochschule Nürnberg(纽伦堡工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出口语维基百科演示语料库,利用LLM生成幻灯片,结合多模态ASR评估,最佳模型微WER达10.23%,并证明跨模态上下文可提升识别性能。
AI 中文摘要
我们提出了口语维基百科演示语料库(Spoken Wikipedia Presentation Corpus),这是口语维基百科语料库(Spoken Wikipedia Corpora)的一个扩展,其中包含由大语言模型(LLM)生成的幻灯片,用于多模态自动语音识别(ASR)。幻灯片由LLM分段的章节通过一个混合流程生成,该流程将基于LLM的内容规划与基于规则的版面设计决策相结合。对于每个章节,LLM生成幻灯片标题、要点、核心信息以及用于创建插图的视觉描述。随后,基于规则的匹配选择布局、主题和样式,以生成最终幻灯片。一个视觉LLM将幻灯片文本提取为Markdown格式。我们评估了多个ASR和口语语言模型(SLMs)。最佳模型在仅音频输入上实现了平均微词错误率(micro-WER)10.23%和平均微字符错误率(micro-CER)6.48%。英语的错误率最低,其次是德语和荷兰语,而在低资源语言上性能有所下降。尽管仅音频的基线模型表现强劲,但对全模态模型(omni models)进行多模态零样本提示仍然具有挑战性。对齐的幻灯片、文本和音频数据显示出通过跨模态上下文改善识别的巨大潜力。
英文摘要
We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.
CommentsAccepted at SLT 2026