arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从高召回到高效用:面向LLM生成客户意图的数据集自适应后处理

From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents

Mahesh Viswanathan, Joan Rossello, Leticia Fernandes, Paul Mutawe

arXiv 2610.09039首次发表:更新:

发表机构

Cisco Systems, Inc.(思科系统公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种数据集自适应后处理架构,将LLM高召回提取的客户意图经标准化、去重、聚类和受约束聚合,转化为高效用、可追踪的意图单元,并应用于下游意图推断与语义检索映射。

AI 中文摘要

大型语言模型可以从异构企业数据中提取有用信号,但高召回率的提取往往会产生重复、粒度不均、语义重叠或数量过多的输出,使得下游系统和人工审核者难以有效使用。我们提出了一种为客户意图提取(CIE)开发的数据集自适应后处理架构,该架构将非结构化的客户语言转化为稳定、可追踪的意图单元。该方法将面向召回率的提取与面向效用的缩减分离。首先,特定于数据源的预处理从多模态计划、稀疏运营记录和结构化机会数据中隔离出证据。随后,候选意图被标准化和去重,可选地通过元数据丰富以用于嵌入计算,在共享的语义向量空间中表示,并根据候选集的特征选择聚类策略进行分组。聚类级关键词提供了可解释性层,而单例重新分配要求嵌入相似性和关键词相似性达成一致。最后,受约束的语言模型聚合为每个聚类生成一个简洁的意图,而不引入不支持的概念,生成的单元保留来源、聚类、嵌入和生成元数据。这将后处理视为一种语义缩减层,而非表面清理,将高召回率的LLM输出转化为可复用的企业智能。我们还描述了两个下游应用:机器生成意图,通过相似画像的同行客户为缺乏直接证据的客户推断可能的目标;以及意图引导的语义检索与映射,将稳定意图作为查询用于下游决策空间,此处以将客户意图映射到业务结果为例进行说明。

英文摘要

Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effectively. We present a dataset-adaptive post-processing architecture developed for Customer Intent Extraction (CIE), where unstructured customer language is transformed into stable, traceable intent units. The approach separates recall-oriented extraction from utility-oriented reduction. Source-specific preprocessing first isolates evidence from multimodal plans, sparse operational records, and structured opportunity data. Candidate intents are then standardized and deduplicated, optionally enriched with metadata for embedding computation, represented in a shared semantic vector space, and grouped using a clustering strategy selected according to the candidate set's characteristics. Cluster-level keywords provide an explainability layer, while singleton reassignment requires agreement between embedding and keyword similarity. Finally, constrained language-model aggregation produces one concise intent per cluster without introducing unsupported concepts, and the resulting unit retains provenance, clustering, embedding, and generation metadata. This treats post-processing not as cosmetic cleanup, but as a semantic reduction layer converting high-recall LLM outputs into reusable enterprise intelligence. We also describe two downstream applications: Machine-Generated Intents, which infer likely objectives for customers lacking direct evidence from peer customers with similar profiles, and intent-guided semantic retrieval and mapping, which uses the stable intent as a query against a downstream decision space, illustrated here by mapping customer intents to business outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑