RAILS:面向大规模场景的检索增强增量式LLM聚类
RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale
浏览论文内容
中文总结 AI 辅助
RAILS通过检索增强和增量式批处理,将LLM聚类扩展到生产规模,在六个基准上平均超越先前最强方法,并成功部署于SaaS工单主题发现场景。
中文摘要 AI 辅助
在生产规模下使用大型语言模型(LLM)作为聚类器面临挑战:提示词无法容纳整个标签空间,且逐文档串行处理无法满足真实工作负载所需的吞吐量。我们提出RAILS,一种检索增强的增量式LLM聚类器,它将聚类转化为对不断增长的标签池的简单循环,并通过文档批处理和有界并发实现扩展。在六个公开基准测试上,RAILS平均超过了先前最强的LLM聚类方法,将准确率从51.2%提升至59.3%,NMI从67.2%提升至74.8%,ARI从45.4%提升至54.7%。我们还报告了来自SaaS工单主题发现管道的生产部署证据,在该管道中,RAILS已取代传统的HDBSCAN阶段,提供了更高的聚类质量、透明的提示驱动控制以及有状态的增量操作。
英文摘要
Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency. On six public benchmarks RAILS exceeds the strongest prior LLM-clustering method on average, lifting accuracy from 51.2% to 59.3%, NMI from 67.2% to 74.8%, and ARI from 45.4% to 54.7%. We further report production-deployment evidence from a SaaS ticket-topic-discovery pipeline, where RAILS has replaced a traditional HDBSCAN stage with higher clustering quality, transparent prompt-driven control, and stateful incremental operation.
发表机构
- Zendesk
机构由 AI 辅助整理,请以论文原文为准。