突破与防御:基于大语言模型的社交媒体机器人检测系统
Breaking and Defending LLM-Powered Social Media Bot Detection Systems
浏览论文内容
中文总结 AI 辅助
本研究针对基于大语言模型的社交媒体机器人检测系统,提出两种新型对抗攻击策略使其准确率降48%,并开发多LLM防御框架LSABRE,在强自适应对抗下仍维持86%准确率,相关方法可推广至其他LLM网络安全系统。
中文摘要 AI 辅助
社交媒体机器人的兴起构成持续威胁,可传播错误信息、操纵舆论并侵蚀用户对在线平台的信任。为应对这一问题,机器学习系统已被开发用于检测和限制机器人活动,但攻击者不断通过对抗学习、行为模仿等技术调整策略,引发机器人与检测工具之间持续的军备竞赛。近期大语言模型(LLM)的进展通过对账号及其内容进行更深入的语义和语境分析,显著提升了机器人检测性能,但这种转变也引入了新的攻击面,使攻击者能够制作直接针对基于LLM分类器的推理与生成机制的漏洞。Anthropic的Claude Code Security等行业工具也利用LLM进行安全关键决策,进一步推动对其攻击面的细致研究。本研究调查了基于LLM的、针对特定威胁的网络安全应用的攻防两方面内容,虽以社交媒体机器人检测挑战为核心,但其方法和见解可推广至基于LLM的广泛网络安全系统,包括钓鱼检测、邮件分类和欺诈分析。我们提出两种新型对抗攻击策略,系统利用基于LLM分类器的语义和语境弱点,使其检测准确率最多下降48%。为应对这些威胁,我们提出一种鲁棒的多LLM防御架构,旨在在自适应对抗条件下保持检测可靠性。我们的解决方案LSABRE(基于大语言模型的社交对抗机器人识别集成)是一个多LLM框架,可大幅提升对一系列攻击的鲁棒性,即使在强大的自适应对抗压力下仍能维持86%的检测准确率。
英文摘要
The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but attackers continuously adapt through techniques such as adversarial learning and behavior imitation, fueling an ongoing arms race between bots and detection tools. Recent advances in large language models (LLMs) have significantly improved bot detection by enabling deeper semantic and contextual analysis of accounts and their content. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target the reasoning and generation mechanisms of LLM-based classifiers. Industry tools such as Anthropic's Claude Code Security similarly leverage LLMs for security-critical decisions, further motivating a careful study of their attack surfaces. In this work, we investigate both the offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, and fraud analysis. We introduce two novel adversarial attack strategies that systematically exploit the semantic and contextual weaknesses of LLM-based classifiers, degrading their detection accuracy by up to 48%. To counter these threats, we propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE (LLM-powered Social Adversarial Bot Recognition Ensemble), is a multi-LLM framework that substantially improves robustness across a range of attacks, maintaining 86% detection accuracy even under strong, adaptive adversarial pressure.
发表机构
- Reichman University(莱希曼大学)
机构由 AI 辅助整理,请以论文原文为准。