arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23075cs.CRcs.AIcs.LG

用于虚假订单欺诈检测的可追溯大语言模型推理

Traceable LLM Reasoning for Fake-Order Fraud Detection

Siqi You, Bingsong Xu, Zhixian Zheng, Xinjian Peng, Yang Xie, Ying Wang, Jiarong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模虚假订单欺诈检测难题,提出基于大语言模型的DeepScrub框架,通过语义统一、持续预训练及SURE机制实现可追溯推理。实验及试点结果显示,该框架提升了欺诈审查精度与召回率,减少工作量并节省成本。

中文摘要 AI 辅助

大规模检测虚假订单欺诈对大型线上到线下(O2O)服务平台来说仍是关键挑战,现有方法依赖专家设计特征、决策不可解释且缺乏推理过程追溯。为此提出基于大语言模型的强化学习框架DeepScrub用于可追溯推理的虚假订单欺诈检测。该框架有三项创新,包括语义统一模块、持续预训练及SURE机制。实验表明,在真实数据集上,DeepScrub宏F1得分达85.3%,优于最佳基线2.7个百分点;在四周现场试点中,精度达91.8%,召回率达88.5%,减少94%的初审工作量,每年节省近百万元,提高了欺诈审查准确性,减少初审工作量并提供可追溯证据。

英文摘要

Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To address these limitations, we propose DeepScrub, a reinforcement learning framework built upon large language models (LLMs) for fake-order fraud detection with traceable reasoning. DeepScrub introduces three innovations. First, a semantic unification module converts heterogeneous risk signals into textual descriptions that LLMs can understand. Second, continued pre-training on risk-control corpora injects domain knowledge, and task rewards jointly evaluate prediction correctness and reasoning quality. Third, the SUggest-REflect (SURE) mechanism incorporates expert feedback and model self-checking to iteratively refine reasoning paths. On a real-world fake-order fraud detection dataset, DeepScrub achieves a macro-F1 score of 85.3%, outperforming the best baseline by 2.7 percentage points. Our task-optimized 8B model further surpasses a 32B model, showing that domain adaptation can matter more than model scale in this setting. In a four-week live pilot, DeepScrub achieved 91.8% precision and 88.5% recall, improving over first-stage human reviewers by 16.6 and 38.8 percentage points. It reduced first-stage manual review workload by 94% and saved nearly one million RMB annually. These results show that DeepScrub improves fraud review accuracy, reduces first-stage review workload, and provides traceable evidence for production risk-review workflows.

发表机构

  • ByteDance(字节跳动)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑