AI 中文总结
该研究提出以事件为中心的AMBER方法,将交互快照压缩为事件Token,在工业级推荐基准上提升了计算-质量帕累托前沿,且事件Token可跨模型架构迁移并随容量扩展持续优化。
AI 中文摘要
基于大语言模型(LLM)的推荐系统已随模型容量和序列长度实现规模化,但每个位置仅编码文本、语义ID或少量类别特征,丢弃了每个事件中可用的丰富用户、物品、上下文及结果信号。在自回归建模下,这导致每个位置的查询能力较弱,且由于每个位置会成为下一个位置的上下文,性能下降会在序列中累积。我们提出一种以事件为中心的范式,用完整时间快照表示每次交互,并确定一个新的可扩展维度,称为快照分辨率:每个事件编码的信息量。为高效扩展快照分辨率,我们引入AMBER(Autoregressive Modeling via Bottlenecked Event Representation,基于瓶颈事件表示的自回归建模),它将每个时间快照压缩为紧凑的事件Token,这是一种新的LLM输入模态。该表示通过端到端学习,同时事件Token可预计算并缓存以用于服务,使快照分辨率与实时服务计算解耦。在工业级排序和检索基准上,AMBER相对于其他推荐范式提升了计算-质量帕累托前沿。在容量充足时,单一统一分词器甚至优于针对各实体类型的专用分词器,证明了在结构不同的实体类型间存在正迁移。AMBER的事件Token还可跨模型架构迁移:当作为服务时的历史特征集成到高度优化的非LLM排序器中时,可产生统计显著的性能提升。进一步扩展事件分词器容量可带来额外性能提升。
英文摘要
LLM-based recommendation has scaled along model capacity and sequence length, yet each position encodes only text, semantic IDs, or a few categorical features, discarding rich user, item, context, and outcome signals available at each event. Under autoregressive modeling, this yields weak queries at each position and, since each position becomes context for the next, the degradation compounds across the sequence. We propose an event-centric paradigm that represents each interaction by its full temporal snapshot, and identify a new scaling dimension we term snapshot resolution: the amount of information encoded per event. To efficiently scale snapshot resolution, we introduce AMBER (Autoregressive Modeling via Bottlenecked Event Representation), which compresses each temporal snapshot into a compact Event Token, a new LLM input modality. The representation is learned end-to-end, while Event Tokens are pre-computed and cached for serving, decoupling snapshot resolution from real-time serving compute. On industrial-scale ranking and retrieval benchmarks, AMBER advances the compute-quality Pareto frontier relative to alternative recommendation paradigms. At sufficient capacity, a single unified tokenizer even outperforms dedicated per-entity tokenizers, demonstrating positive transfer across structurally different entity types. AMBER's Event Tokens also transfer across model architectures: when integrated into a heavily optimized non-LLM ranker as serving-time historical features, they yield statistically significant improvements. Further scaling Event Tokenizer capacity provides additional improvements.
Comments11 pages, 10 figures, 7 tables