AI 中文总结
SeqLLM是一种在保留LLM语言能力的同时增强其行为序列建模能力的框架,已部署于微信支付,在风险筛查等任务上性能显著优于现有基线。
AI 中文摘要
大型支付平台的商户风险控制每日需要筛查数千万商户,其中误判会损害合法商户,而漏判则会让有害活动未被发现。最棘手的案例需要同时理解商户的文本资料和长期行为序列。大语言模型(LLM)擅长处理文本,但无法原生地对这类序列进行建模,而对其进行适配往往会导致灾难性遗忘。我们提出了SeqLLM,这是一种在保留预训练LLM语言能力的同时为其添加行为序列建模能力的框架。SeqLLM包含三个组件:一是将行为事件表示为原生token的紧凑离散词汇表;二是通过两阶段对齐课程训练的轻量投影器,将这些token映射到LLM的语义空间中;三是前缀引导的能力注入,该组件通过任务前缀的监督微调而非持续预训练来获取序列建模能力。SeqLLM已部署在微信支付,每日筛查数百万商户。与基于DeepSeek的生产环境LLM基线相比,它将筛查精度从92.0%提升至97.5%。其预训练的行为token嵌入还将服务于十亿级交易流量的生产环境欺诈检测器的Top-0.01%精度提升了26.8个百分点。除支付领域外,SeqLLM在公开推荐基准上也取得了最先进的结果。在MovieLens和Amazon数据集上,它超越了强大的User-LLM基线,在Recall@5指标上相对提升最高达32%,同时保留了明显更强的语言能力。在RecIF基准上,它仅用OneRec-8B流水线五分之一的GPU天数,就将Pass@32指标提升了14.2%。
英文摘要
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly understanding a merchant's textual profile and long behavioral sequence. Large language models (LLMs) excel at text but cannot natively model such sequences, while adapting them often causes catastrophic forgetting. We present SeqLLM, a framework that adds behavioral-sequence modeling to a pretrained LLM while preserving its language ability. SeqLLM combines three components: a compact discrete vocabulary that represents behavioral events as native tokens; a lightweight projector, trained with a two-stage alignment curriculum, that grounds these tokens in the LLM's semantic space; and prefix-guided capability injection, which acquires sequence-modeling ability through task-prefixed supervised fine-tuning rather than continual pre-training. SeqLLM is deployed at WeChat Pay, screening millions of merchants daily. Against the production DeepSeek-based LLM baseline, it raises screening precision from 92.0% to 97.5%. Its pretrained behavior-token embeddings also improve Precision@Top-0.01% by 26.8 percentage points in a production fraud detector serving billion-scale transaction traffic. Beyond payments, SeqLLM achieves state-of-the-art results on public recommendation benchmarks. On MovieLens and Amazon, it surpasses the strong User-LLM baseline by up to 32% relative Recall@5 while retaining markedly stronger language ability. On RecIF, it improves Pass@32 by 14.2% over the full OneRec-8B pipeline using only one-fifth of its GPU-days.