发表机构
Amity Research and Application Center (ARAC), Amity AI Holdings Co., Ltd.(友好研究与应用中心(ARAC),友好人工智能控股有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在解决客户行为建模问题,提出大型行为模型(LBM),通过统一公式从零售交易学习客户决策,经多种训练方式优化。该模型在多任务中表现出色,消融研究揭示各因素作用,为客户数字孪生体和行为模拟提供可扩展基础。
AI 中文摘要
客户行为建模是推荐、营销和决策支持的基础,但现有方法要么优化预测准确性却不解释决策,要么模拟用户却不以真实行为数据为依据。我们提出大型行为模型(LBM),它通过统一的人-环境公式直接从大规模零售交易中学习客户决策。客户状态由从历史购买中得出的行为概况表示,产品上下文通过检索增强生成纳入。该模型通过对语言化行为数据进行持续预训练、对决策生成进行监督微调以及使用可验证奖励进行强化学习以进行基于证据的校准来训练。我们在购买预测、硬负例判别、购物篮完成、促销响应和跨域优惠券兑换等方面评估了该框架。该模型在零售领域任务上始终优于前沿通用语言模型,同时在跨零售商和决策领域展示出强大的零样本和微调迁移能力。消融研究表明持续预训练是行为泛化的主要驱动因素,检索在训练和推理期间应用时最有效,强化学习提高了对明确行为证据的依赖。这些结果表明交易历史中编码的行为知识可以被语言模型有效学习,为客户数字孪生体和行为模拟提供了可扩展基础。
英文摘要
Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation. Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation. The model is trained using continued pre-training on verbalized behavioral data, supervised fine-tuning for decision generation, and reinforcement learning with verifiable rewards for evidence-based calibration. We evaluate the proposed framework on purchase prediction, hard-negative discrimination, basket completion, promotion response, and cross-domain voucher redemption. The model consistently outperforms frontier general-purpose language models on in-domain retail tasks while demonstrating strong zero-shot and fine-tuned transfer across retailers and decision domains. Ablation studies show that continued pre-training is the primary driver of behavioral generalization, retrieval is most effective when applied during both training and inference, and reinforcement learning improves reliance on explicit behavioral evidence over generic language-model priors. These results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.
Comments17 pages, 5 figures