复用潜在语音表示用于转录文本中查询条件主题定位
Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts
浏览论文内容
中文总结 AI 辅助
本文研究查询条件主题定位,通过复用ASR编码器状态与文本嵌入融合,实现无需额外音频编码器的轻量级跨度定位,在公开数据集上超越纯文本基线,对结构化语音效果最佳。
中文摘要 AI 辅助
长转录文本对于下游NLP系统而言是昂贵的输入,且通常包含不相关上下文。我们研究查询条件主题定位:预测转录文本中最能回答主题标题查询的句子跨度。为改进跨度定位,我们复用ASR编码器状态作为句子级表示,并将其与文本嵌入融合。这使得轻量级跨度定位器无需运行单独的音频编码器即可利用语音信息。在两个公开数据集上的实验表明,与纯文本基线相比,该方法持续取得改进,尤其在严格边界匹配标准下。跨数据集实验进一步表明,对于结构化或半结构化语音,收益最为显著,而对自发语音的收益有限且不一致。
英文摘要
Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.
发表机构
- Technische Hochschule Nürnberg Georg Simon Ohm(纽伦堡应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。