HubMixer:用于推荐系统中参数高效特征交互的渐进式潜在中心混合
HubMixer: Progressive Latent Hub Mixing for Parameter-Efficient Feature Interaction in Recommendation
AI总结:
本文提出 HubMixer,一种参数高效的潜在中心混合架构,通过“诱导—交互—读出”范式实现推荐系统的特征交互,经离线实验优于 SOTA 模型,在线 A/B 测试使快手招聘简历提交转化率提升 5.48% 并已落地部署。
AI中文摘要:
学习有效的特征交互是工业推荐和广告排序系统的核心。近期的 token 混合架构用轻量混合算子简化了自注意力,提升了硬件效率并支持大规模部署。然而,推荐系统的 token 本质上是异质的:用户画像、物品属性、行为序列、上下文特征、统计信号和业务侧特征属于不同语义空间,且以稀疏、样本特定的模式交互。因此,在原始异质 token 空间中直接混合所有 token 可能参数效率低下,因为模型必须隐式发现哪些特征组应交互以及如何路由此类交互。本文提出 HubMixer,一种用于推荐系统特征交互的参数高效潜在中心混合架构。HubMixer 不直接混合原始特征 token,而是引入少量可学习的潜在中心,通过“诱导—交互—读出”范式组织特征交互:首先,中心诱导将异质 token 汇总为紧凑的潜在中心,潜在中心通过交叉注意力查询输入 token;其次,中心交互在更干净的潜在中心空间中执行高阶交互;最后,token 条件读出让每个原始 token 选择性地从交互后的中心读取,注入全局交互语义同时保留 token 级域身份。在工业推荐任务上的大量离线实验表明,HubMixer 优于 SOTA 模型;在快手短视频招聘业务的在线 A/B 测试进一步显示,简历提交转化率实现了具有统计显著性的 5.48% 提升,且 HubMixer 已完全部署到生产环境。
英文摘要:
Learning effective feature interactions is central to industrial recommendation and advertising ranking systems. Recent token-mixing architectures simplify self-attention with lightweight mixing operators, improving hardware efficiency and enabling large-scale deployment. However, recommendation tokens are fundamentally heterogeneous: user profiles, item attributes, behavioral sequences, context features, statistical signals, and business-side features live in different semantic spaces and interact in sparse, sample-specific patterns. Directly mixing all tokens in the raw heterogeneous token space may therefore be parameter-inefficient, as the model must implicitly discover which feature groups should interact and how such interactions should be routed. In the paper, we propose HubMixer, a parameter-efficient latent hub mixing architecture for feature interaction in recommendation. Instead of directly mixing raw feature tokens, HubMixer introduces a small set of learnable latent hubs to organize feature interactions through an `induction--interaction--readout` paradigm. First, hub induction summarizes heterogeneous tokens into compact latent hubs, where latent hubs query input tokens through cross-attention. Second, hub interaction performs high-order interaction in the cleaner latent hub space. Third, token-conditioned readout lets each original token selectively read from the interacted hubs, injecting global interaction semantics while preserving token-level field identity. Extensive offline experiments on industrial recommendation tasks show that HubMixer outperforms the SOTA models. Online A/B testing in the Kuaishou short-video recruitment business further shows a statistically significant 5.48% improvement in resume submission conversion rate, and HubMixer has been fully deployed in production.