arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型在欺诈检测和信任与安全工作流程中的操作证据差距

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

Keyur Gabani

arXiv 2607.13078首次发表:更新:

AI 中文总结

研究LLMs用于欺诈检测等信任与安全工作流程时的操作证据差距。通过对相关操作源编码调查发现证据失衡,欺诈任务部分大但缺乏关键证据。贡献了FORTE框架和部署证据清单,确定支持LLMs部署所需研究。

AI 中文摘要

大语言模型(LLMs)如今被用于欺诈检测、诈骗调查、内容审核及其他信任与安全工作流程。多数公开文献仍将其作为模型评估,较少关注其在实际操作流程中的行为。这引发一个实际证据问题:在有延迟、成本、升级、人工审核和对抗风险限制的实时工作流程中,何种依据能证明将LLM置于其中是合理的?我们通过对部署证据进行以欺诈为先的调查来解决此问题。我们对49个与LLM在欺诈检测、调查支持、内容审核及交叉稳健性方面使用相关的操作源进行编码(18个欺诈、14个审核、17个交叉),并辅以15个确定调查边界的背景参考。这些源包括系统、基准、框架及与部署相关的调查,而非49个生产部署。主要发现是证据失衡。欺诈在编码语料库中占最大的特定任务部分。然而,审核论文包含更多关于延迟、成本、治理和公平性的明确公开证据。在18个欺诈和调查源中,没有一个报告清晰的每次决策延迟、每次决策成本或校准证据;大多数报告的是离线任务性能、检索增益或案例研究准确性。该调查贡献了一个角色和证据组织框架FORTE,用于将LLM定位为分类器、检索接口、解释生成器、审核助手、代理、特征提取器或升级组件。它还贡献了一个最小部署证据清单,涵盖延迟预算、每次决策成本、决策阈值、解释完整性和对抗压力。由此产生的议程确定了支持基于LLM的欺诈和信任与安全工作部署主张所需的研究。

英文摘要

LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.

Comments22 pages, 3 figures, 6 tables. Ancillary files include the evidence matrix, search note, and numeric claim check

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑