发表机构
J.P. Morgan Chase(摩根大通)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SeLATM框架,通过段级主题生成和智能体反馈循环细化主题,在减少LLM资源消耗的同时保持优越性能,以解决主题分配过程的缺陷。
AI 中文摘要
主题建模是一种在文档中发现隐藏主题的有效技术,广泛应用于各行业领域的文本挖掘和数据分析中。近年来,基于大型语言模型(LLM)的主题模型应运而生,它们通过提示LLM生成主题,然后将主题分配给文档,从而产生比传统主题建模算法更自然、更易读的主题。然而,主题分配过程的特性导致了一些缺点,例如无法生成文档上的主题分布、主题过于宽泛或狭窄,以及高资源消耗,且资源消耗随被分配主题的文档数量和长度增加而增长。这些问题对于需要高质量、深入分析以及处理大量文档的工业应用尤为关键。在此背景下,本文提出了一种名为SeLATM的框架,通过采用段级主题生成和基于智能体反馈循环的主题细化来解决上述问题。在多个数据集上的实验结果表明,与基于主题分配过程的方法相比,SeLATM显著减少了LLM资源消耗,同时保持了优越的性能。
英文摘要
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.