arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向十亿级容量用户表示学习的密度定律

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang, Letian Gong, Baokun Wang, Huan Li, Yu Cheng, Weiqiang Wang

arXiv 2608.23392首次发表:更新:

发表机构

DeepFind Team, Ant Group; Zhejiang University(蚂蚁集团DeepFind团队; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对十亿级用户表示学习的原始数据扩展瓶颈与分词配置定量分析缺失问题,提出用户行为密度定律,开发自适应可变分词方法ALGN,经多源实验验证其通用性与可靠性且性能优于基线。

AI 中文摘要

现实工业场景中的用户表示学习通常通过增加用户数量、行为序列长度和模型规模来扩展。然而,现有方法面临两大挑战:(i)十亿级容量下原始数据扩展的瓶颈,即随着原始文本用户行为输入规模增大,性能增益逐渐递减,这一问题可通过分词缓解;(ii)缺乏对分词配置应如何随数据规模扩展的定量分析。本报告提出用户行为密度定律,用于表征数据规模与最小充分分词容量之间的定量关系。首先,我们在十亿级支付宝数据集上开展原始数据与分词数据的扩展对比预实验,揭示了原始数据扩展瓶颈以及分词带来的持续性能增益。为推导不同数据规模下最小充分分词配置的扩展规律,我们结合理论分析与系统实验总结出定量扩展模式。研究发现,最小充分分词容量的对数与以词元(token)衡量的输入数据规模之间存在近似线性关系,且扩展斜率会随分词方法和数据源系统变化,反映出表示空间冗余度和源内唯一性的差异。基于所提出的定律,我们进一步开发了自适应可变分词方法ALGN,该方法可优化容量分配。在不同数据源、分词方法和下游任务上开展的大量实验证明了用户行为密度定律的通用性和可靠性,为大规模用户表示学习中的分词配置选择提供了实用指导,且ALGN的性能优于现有基线方法。

英文摘要

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative relationship between data scale and the minimum sufficient tokenization capacity. Firstly, we conduct a pilot study on raw & tokenized scaling comparison on billion-scale Alipay dataset, revealing the raw data scaling bottleneck and the sustained gains enabled by tokenization. To derive the scaling pattern governing the minimum sufficient tokenization configuration at different data scales, theoretical analysis and systematic experiments are employed to summarize the quantitative scaling pattern. We find an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size measured by tokens, and the scaling slope varies systematically with the tokenization method and data source, reflecting differences in representation-space redundancy and intra-source uniqueness. Guided by the proposed law, we further develop ALGN, an adaptive variable-length tokenization method that improves capacity allocation. Extensive experiments across diverse data sources, tokenization methods, and downstream tasks demonstrate the generalizability and reliability of the User Behavioral Densing Law, providing practical guidance for tokenization configuration selection in large-scale user representation learning. Moreover, ALGN outperforms existing baselines.

Comments28 pages, 13 figures, technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑