arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11707cs.MAcs.AIcs.CRcs.HC

基于摘要记忆的可解释智能系统用于检测对话诈骗

An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory

Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare

首次发表
浏览论文内容

中文总结 AI 辅助

针对对话诈骗威胁,本文提出基于摘要记忆的可解释智能系统,扩展单消息钓鱼检测。引入多类别基准ConScamBench - 278,单消息和对话级检测器效果良好,两项用户研究及可用性验证了系统有效性。

中文摘要 AI 辅助

随着生成式人工智能的快速发展,对话诈骗带来的威胁日益增长。现有诈骗检测系统主要关注孤立消息,不足以应对这种不断演变的威胁。本文扩展了单消息网络钓鱼检测,提出了一个用于检测复杂对话诈骗的可解释智能系统,并引入了ConScamBench - 278这一用于对话诈骗检测的多类别基准。单消息检测器在孤立消息上实现了100%的网络钓鱼召回率,对话级检测器在公共LoveFraud02语料库中识别出所有对话诈骗,在ConScamBench - 278上达到97.8%的准确率。两项用户研究进一步验证了该系统,系统的可用性量表得分也高于既定基准。

英文摘要

Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated conversational scams. It also introduces ConScamBench-278, an initial public multi-category benchmark for conversational scam detection spanning eight scam types, released to support reproducible evaluation and future expansion. On isolated messages the single-message detector attains 100% phishing recall, while the conversation-level detector identifies all conversational scams in the public LoveFraud02 corpus (83/83) and reaches 97.8% accuracy (95% CI [95.4, 99.0]) on ConScamBench-278. Two user studies (N = 100 and N = 45) further motivate the system: participants report frequently experiencing uncertainty when judging suspicious conversations. In an uncontrolled pre/post comparison, users self-reported trust, self-confidence, and perceived need for AI-based scam detection all increased (p < 0.001, Wilcoxon signed-rank). The system also receives a System Usability Scale score of 74.7 (95% CI [72.5, 76.9]), above the established usability benchmark.

发表机构

  • University of New Haven(新罕布什尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑