arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大规模LLM推理中具有可证明护栏的高效聚类

Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale

Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer

arXiv 2607.19704首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大规模LLM推理成本和延迟瓶颈,提出两阶段聚类算法,先用Mini-batch K-Means生成初始聚类,再在其中选代表,能保证逐样本质量控制,运行高效且可扩展,相比常见方法有显著优势,部署后大幅降低成本和延迟。

AI 中文摘要

将基于大语言模型(LLM)的应用扩展到数百万用户,受到现代基础模型推理成本和延迟的瓶颈限制。一种自然的解决方法是对输入进行聚类,仅对聚类代表调用LLM,让其他成员继承输出,但这仅在每个成员与其代表可测量地接近时才安全。现有聚类方法在大规模时无法提供这种逐样本质量控制。我们提出一种两阶段算法,首先用Mini-batch K-Means生成初始聚类,然后在每个初始聚类中贪婪地选择代表。该算法通过构造精确地强制相似性和属性护栏,对于n个样本、特征维度d和K个初始聚类,运行时间为$O(nd + n^2 d/K)$,内存为$O(nd + n^2/K^2)$,当K与n成比例增长时,在n上呈线性。我们在内部和公共数据集上与常见聚类方法进行了基准测试,我们的方法不仅提供逐样本护栏,而且运行速度快10 - 1000倍,能够扩展到大多数标准方法难以处理的数据大小。在基于人物角色的推荐系统中部署在3800万客户上时,该聚类方法将下游成本和延迟降低了50倍,同时保留了个性化并推动了生产发布。

英文摘要

LLMs can be prohibitively expensive and slow to run at scale, especially for applications that invoke an LLM per sample over millions of inputs. A natural way to scale is to cluster the inputs, run the LLM only on cluster representatives, and propagate the outputs to other cluster members. However, the outputs a member receives are only as good as its match to the representative. Off-the-shelf clustering methods optimize an aggregate objective, targeting average-case quality without per-sample guardrails. As a result, members can be assigned to poorly-matched representatives, and the inherited outputs -- though appropriate for the representative -- may be irrelevant or even unsafe for the member. For example, a parent of a toddler grouped with parents of older children could receive age-inappropriate recommendations. Most clustering methods also scale poorly to millions of inputs in runtime and memory, limiting their use at industry scale. We propose a scalable two-stage clustering algorithm with provable per-sample guardrails: every sample is guaranteed to share a user-specified minimal embedding similarity and exact attribute match with its representative. The algorithm first generates initial clusters with Mini-batch K-Means, then greedily selects representatives within each to satisfy the guardrails. We provide theoretical guarantees, complexity analysis, and benchmarks against common methods on internal and public datasets. Our method delivers per-sample guardrails while running substantially faster and scaling to data sizes where most standard methods become intractable. We demonstrate its impact in a real-world deployment clustering 38 million customers, reducing downstream LLM cost and runtime by 50-fold while preserving personalization. This unblocked the launch of a persona-based recommender system that delivers significant gains in revenue and engagement in an A/B test.

CommentsV1 accepted at High Dimensional Learning Dynamics workshop, ICML 2026 (non-archival), V2 accepted at The Third Workshop on Agentic and Generative AI for E-Commerce, RecSys 2026 (non-archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑