arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34425cs.CLcs.AI

口语文档的零样本线索引导主题分割

Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents

  • Seoul National University(首尔大学)
  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Suhwan Choi, Myeongho Jeon, Myungjoo Kang

AI总结:

针对口语文档主题分割中粒度适应性差的问题,提出无需训练的线索引导分割框架CGS,先利用显式主题转换短语定位边界,线索不足时回退至语义分割,在六个基准和六个LLM上持续超越基线且对ASR噪声鲁棒、API成本低。

AI中文摘要:

主题分割将口语文档组织成连贯的章节,便于导航和下游理解。合适的粒度可能有很大差异,从宽泛的主题转变到细粒度的子主题。然而,现有的基于LLM的分割器往往难以适应这种变化,导致它们要么合并不同的子主题,要么过度分割连贯的主题。为了解决这个问题,我们引入了线索引导分割(Cue-Grounded Segmentation, CGS),这是一种无需训练且不依赖任何任务特定监督的框架。CGS首先识别明确指示新主题开始的短语,并将其句子位置用作分割边界。当此类线索不足时,它会回退到语义分割,由线索提取过程中推断出的文档结构引导。在六个基准测试和六个LLM骨干网络上,CGS持续优于现有基线,对嘈杂的ASR转录本保持稳健,并在专有模型上以较低的API成本实现这些增益。

英文摘要:

Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.

↑