arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

野外场景下的主题匹配:来自真实世界ASR转录文本的基准与经验教训

Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw

arXiv 2609.00330首次发表:更新:

发表机构

Dialpad Inc.(戴尔德公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对联络中心ASR转录的主题匹配任务,构建了人工标注数据集,对比了三类匹配器及两种主题表示,发现轻量级Gemini LLM匹配器搭配自然语言描述时性能最优。

AI 中文摘要

在联络中心,实时智能体辅助工具需针对多个预定义主题,判断实时客户话语是否相关,并在相关时向智能体显示指导卡片。输入存在噪声且极具挑战性:自动语音识别(ASR)转录的自发电话对话,内容可能不清晰、重复,且大多缺乏标点。为系统研究这一现实任务,我们整理了源自真实呼叫中心转录文本的人工标注主题-话语判断数据集。我们比较了三类匹配器:基于正则表达式的基线、零样本句子嵌入编码器、基于Gemini的大语言模型(LLM)匹配器。此外,我们在基准中研究了两类主题表示:关键词和自然语言描述。我们的实证实验表明,当配备自然语言描述时,轻量级LLM匹配器的性能优于嵌入模型和正则表达式模型。

英文摘要

In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speech Recognition) transcripts of spontaneous phone conversations, which can be unclear, repetitive, and mostly lack punctuation. To systematically study this real-world task, we curate a human-annotated topic-utterance judgments dataset sourced from real call-center transcripts. We compare three types of matchers: a regex-based baseline, zero-shot sentence embedding encoders, and Gemini-based LLM matchers. In addition, two types of topic representations are studied in our benchmark:keyphrases and natural language description. Our empirical experiments highlight the superior performance of lightweight LLM matchers over embedding and regex models when equipped with natural language descriptions.

CommentsAccepted at the 11th Workshop on Natural User-generated Text (W-NUT 2026), EMNLP 2026. Camera-ready version. 9 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑