发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模非结构化反馈文本分类法构建的挑战,本文提出TaxCE框架,结合EEG指标的闭环迭代优化机制,在分类法质量上优于现有基线方法。
AI 中文摘要
将非结构化反馈文本组织为层次化分类法是自然语言处理(NLP)领域的基础挑战,尤其适用于反馈以评论、转录文本、调查等多样形式大规模涌现的场景。现有方法要么生成的层次结构较浅,要么忽略长尾主题,要么缺乏严谨的评估框架。本文提出TaxCE,这是一个全自动化框架,通过逐步将语料库内容浓缩为可操作片段、去重语义单元及带定义的细粒度主题,再自下而上组织为具备语料库依据的层次结构,以此从原始文本构建多级层次化分类法。本文还引入三项基于语料库的评估指标:排他性(Exclusivity)、穷尽性(Exhaustivity)和粒度(Granularity,简称EEG),并将其整合至指标闭环迭代优化机制中,该机制可诊断缺陷并应用针对性修正,直至收敛。大量实验表明,TaxCE在涵盖经典主题模型、神经方法及基于大语言模型(LLM)的方法的现有基线中表现始终更优,相较于最强基线,其在排他性、穷尽性和粒度上的平均提升分别为11.8、20.5和15.7个百分点。人工评估进一步证实TaxCE生成的分类法在质量、可操作性和可导航性上更优。
英文摘要
Organizing unstructured feedback text into hierarchical taxonomy is a fundamental challenge in NLP, particularly in domains where feedback arrives at massive scale in varied forms such as reviews, transcripts, and surveys. Existing approaches either produce shallow hierarchies, neglect long-tail topics, or lack rigorous evaluation frameworks. We present TaxCE, a fully automated framework that constructs multi-level hierarchical taxonomies from raw text through progressive condensation of corpus content into actionable segments, deduplicated semantic units, and granular topics with definitions, which are then organized bottom-up into a hierarchy with corpus-groundedness. We also introduce three corpus-grounded evaluation metrics, Exclusivity, Exhaustivity, and Granularity (EEG), and integrate them into a metrics-in-the-loop iterative refinement mechanism that diagnoses deficiencies and applies targeted corrections until convergence. Extensive experiments demonstrate that TaxCE consistently outperforms existing baselines spanning classical topic models, neural methods, and LLM-based approaches, with average improvements of 11.8, 20.5, and 15.7 percentage points in exclusivity, exhaustivity, and granularity respectively over the strongest baseline. Human evaluation further confirms superior taxonomy quality, actionability, and navigability.
CommentsEMNLP 2026 - Industry