arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10239cs.AI

超越检测:评估防御型大语言模型(LLM)在实时逐轮交互中对抗AI生成社会工程攻击的能力

Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction

Yuqiao Xu, Osama Zafar, Alexander Nemecek, Erman Ayday

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建含300例的在线住房语料库,评估5种防御型LLM在实时逐轮与静态设置下对抗AI社会工程攻击的能力,发现防御效果差异大,需单独测量干预等指标。

中文摘要 AI 辅助

生成式AI使社会工程攻击更流畅、自适应且可扩展,因此亟需基于大语言模型(LLM)的防御工具,在持续交互中保护用户。本研究探究这类防御工具是识别风险的结构性根源,还是仅对表面线索做出反应。我们将信任链定位形式化,即识别交互失败是否源于行为者权限、资产控制、验证充分性或交易路径。我们构建了包含20类场景、合法案例、4种结构性失败模式及3种表面条件的受控300例在线住房语料库。在有状态逐轮交互与一次性静态两种设置下,对5种防御模型在同一语料库上进行评估,每种设置各产生1500次模型-案例评估,总计3000次。没有模型出现明确的不安全合规行为,但防御效果差异显著:干预率介于0%至96.3%之间。保护行动与正确的结构性定位常相互脱节,模型有时会在识别错误信任组件时进行干预,或在识别结构性失败时未采取保护行动。资产控制失败是主要的定位瓶颈,不同模型的表面敏感性存在差异,实时与静态表现的差异因模型而异。这些发现表明,仅安全外观行为是不够的;实时反诈骗需单独测量干预、时机、结构性定位及假阳性行为。

英文摘要

Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions. We ask whether such defenders identify the structural source of risk or merely react to surface cues. We formalize trust-chain localization: identifying whether an interaction fails at actor authority, asset control, verification sufficiency, or transaction path. We construct a controlled 300-case online-housing corpus spanning 20 scenario families, legitimate cases, four structural failure modes, and three surface conditions. Five defender models are evaluated on the same corpus in state- ful turn-by-turn and one-shot static settings, yielding 1,500 model-case evaluations per protocol and 3,000 in total. No model produced explicit unsafe compliance, yet defensive effectiveness varied sharply: intervention rates ranged from 0% to 96.3%. Protective action and correct structural localization were frequently decoupled, with models sometimes intervening while identifying the wrong trust component or recognizing a structural failure without taking protective action. Asset-control failures were a major localization bottleneck, surface sensitivity varied across models, and live-static differences were model-dependent. These findings show that safe-looking behavior alone is insufficient; live scam resistance must separately measure intervention, timing, structural localization, and false-positive behavior.

↑