arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

组件基准:面向大规模推荐系统的层次化模型剖析

Component Benchmark: Hierarchical Model Profiling for Large-scale Recommendation Systems

Dharak Kharod, Yuzhen Huang, Zhou Wang, Jackie Xu, Fuzail Khan, Jacky Zhou, Hao Yan, Lidong Zhao, Xizhou Feng, Yvonne Liu, Karthik Jayaraman, Praveen Ramachandran, Vishwa Karia, Yashasvi Makin

arXiv 2609.30656首次发表:更新:

发表机构

Meta Platforms Inc(Meta平台公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大规模推荐模型结构异构、剖析困难的问题,提出组件基准CB系统,以层次化方式独立剖析各子模块性能并提供树状交互可视化,已在TB级模型上验证有效性。

AI 中文摘要

大规模推荐模型提出了独特且尚未充分探索的剖析挑战。大多数推荐模型架构在结构上是异构的,混合了内存带宽受限的操作、小型计算受限的稠密层、由不规则类别特征导致的动态形状,以及低算术强度操作。推荐模型随着建模工程师尝试各种组合而快速演进,这些组合往往在缺乏对硬件执行特性可见性的情况下编写。标准剖析工具要么提供端到端吞吐量,要么提供算子级跟踪,但无法将性能归因于实践者所推理的子模块。我们提出了组件基准(Component Benchmark, CB),一个剖析系统,它以层次化方式独立表征每个子模块的性能,提供树状结构的交互式可视化,为机器学习实践者带来性能清晰度。其核心是,CB提供了一个简单但可扩展的、基于子模块的基准测试框架,具有插件架构,支持层次化性能分析。这些大规模推荐模型达到TB级规模,在数千个GPU上运行,每天处理1000亿个样本。我们展示了CB在常见开源模型上的有效性,并讨论了CB如何被用于加速现代推荐模型的性能分析和优化。

英文摘要

Large-scale recommendation models pose distinct, under-explored profiling challenges. Most recommendation model architectures are structurally heterogeneous, intermixing memory-bandwidth-bound operations, small compute-bound dense layers, dynamic shapes from jagged categorical features, and low-arithmetic-intensity operations. Recommendation models evolve rapidly as modeling engineers experiment with compositions, often written without visibility into hardware execution characteristics. Standard profiling tools offer either end-to-end throughput or operator-level traces, but cannot attribute performance to the submodules that practitioners reason about. We present Component Benchmark (CB), a profiling system that independently characterizes each submodule performance in a hierarchical manner, providing a tree-structured, interactive visualization that brings performance clarity to ML practitioners. At its core, CB provides a simple yet extensible, submodule-based benchmarking framework with a plugin architecture that enables hierarchical performance analysis. These large-scale recommendation models are TB-scale, run on thousands of GPUs and ingest 100B examples per day. We demonstrate CB's effectiveness on common open sourced models and discuss how CB has been leveraged to accelerate modern recommendation model performance analysis and optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑