先澄清,再聚焦:面向大规模对话分析的陈述规范化
Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
浏览论文内容
中文总结 AI 辅助
该研究提出陈述规范化方法,将对话转换为带引用和标签的说话人陈述,在客服报价抑制任务中提升模型性能,构建小型模型推理流水线以降低大规模对话分析成本。
中文摘要 AI 辅助
企业对话分析需要对数百万次交互提出诸多问题,每个问题都需重构人们的意图并识别关键信息,在相同 transcript(对话记录)上重复开展成本高昂的解释工作。我们提出一条简单原则:先澄清文本,再聚焦读者。陈述规范化将对话转换为带有来源引用和语义标签的、由说话人标注的简短陈述,这些陈述使含义更明确,标签则支持为特定问题选择证据。下游模型可根据决策需求使用完整表示或相关子集。在客服通话的报价抑制任务中,规范化提升了无选择的监督分类器,而较弱的提示式读者则同时受益于规范化和选择。小型模型可学习规范化约定,轻量编码器则处理标签和下游决策。在多个问题间共享该准备工作,可构建完全由小型模型组成的推理流水线,大幅降低对数百万次对话进行分析的成本。
英文摘要
Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.