论提高基于文档的播客忠实度
On Improving Faithfulness of Podcasts from Documents
浏览论文内容
中文总结 AI 辅助
研究基于文档的播客生成中的忠实度问题,构建数据集并用多个大语言模型生成记录,引入轮次级评判框架。发现先进模型也常生成无根据内容,提出catch - n - repair框架,实验证明该框架能提升播客生成的忠实度。
中文摘要 AI 辅助
大语言模型越来越多地用于从文本来源生成如播客这样的长篇对话内容。虽然这些系统能产出流畅且引人入胜的叙述,但常引入无根据信息。本文首次对基于文档的播客生成中的忠实度进行系统研究,需在长篇多说话者文字记录的对话轮次中保持依据。构建超1500个跨五个领域文档的数据集,用多个大语言模型生成播客文字记录,引入轮次级大语言模型评判框架评估对话轮次是否有文档依据,并通过人工研究验证其可靠性。分析表明包括GPT - 4o在内的先进模型也常生成无根据内容。为此提出catch - n - repair框架,能检测并改写不忠实对话轮次同时保留对话流程。实验证明在领域内和领域外设置中忠实度都有持续提升。
英文摘要
Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these systems produce fluent and engaging narratives, they often introduce ungrounded information. In this work, we present the first systematic study of faithfulness in document-grounded podcast generation, where grounding must be maintained across conversational turns in long-form, multi-speaker transcripts. We construct a dataset of over 1500 documents spanning five domains and generate podcast transcripts using multiple LLMs. We introduce a turn-level LLM-as-a-judge framework for evaluating whether conversational turns are supported by the source document, and validate its reliability through human studies. Our analysis shows that even state-of-the-art models, including GPT-4o, frequently generate ungrounded content. To mitigate this issue, we propose catch-n-repair, a model-agnostic framework that detects and rewrites unfaithful conversational turns while preserving conversational flow. Experiments demonstrate consistent improvements in faithfulness across both in-domain and out-of-domain settings.
发表机构
- Indian Institute of Science(印度科学研究所)
- Adobe Research(Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。