arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31045cs.IRcs.CLcs.LG

KuaFu: 将长用户行为压缩为十亿规模的理解

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang, Xining Ran, Ben Tan, Yeshou Cai, Gong Chen, Haijie Gu, Jie Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

KuaFu是一个统一的行为压缩层,通过双轴投影器将每个行为项压缩为2-4个token,采用保真导向训练,在四个生产任务中达到或超过未压缩模型,提升吞吐量37%-350%,节省190个GPU,并在腾讯平台运行十个月提升GMV 1.37%。

中文摘要 AI 辅助

对话代理、生成式推荐系统和个性化广告都依赖于一种能力:从原始行为中理解每个用户。当前的工业实践是任务特定的:对于每个任务,从完整历史中提取相关子序列,并训练专用模型。在生产中,这遇到两个瓶颈。首先,即使经过过滤,单任务序列仍然非常长:内容兴趣摘要每个用户读取数百个项目,序列化为提示文本后可达数万token。其次,用户画像定期刷新:每周十亿用户,总计约10万QPM,在固定GPU预算下,这设定了硬吞吐量下限。因此,压缩是必须的,但截断或粗粒度压缩可能静默地扭曲用户画像,引入四种幻觉类型(捏造、遗漏、日期错误归属、逻辑断裂),由于无法评估压缩表示本身,这些幻觉只能在下游指标中表现为模糊的退化。我们提出KuaFu,一个统一的用户行为压缩层,其最小单元是一个行为项。一个双轴投影器将每个项压缩为2-4个token,宽度128-256(沿token轴约10倍,沿宽度约20倍;每项缓存从10KB降至0.5KB),并采用保真导向的四阶段训练和分层中间评估。在四个生产画像任务中,它在所有五个主要指标上达到或超过未压缩的单任务生产模型,将每GPU吞吐量提高37%-350%,并节省190个GPU。在公共基准上,它在相同压缩比下几乎总是优于先前的压缩器(在域外MRQA上最高提升+17.7 EM);在RecBench上,一个4B模型比其8B对应模型高出1.90分。KuaFu已在腾讯广告和推荐平台运行十个月,整体GMV提升1.37%。

英文摘要

Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.

发表机构

  • Tencent Inc.(腾讯公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑