arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构建它、攻破它、重复它:基准测试与改进大语言模型生成的社交媒体虚假信息检测

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

Kevin Thomas, Milosz Kasprzyk, Reuel C Igbokwe Onuigbo, Elliott Pert, Cameron Tovey, João A. Leite, Olesya Razuvayevskaya, Carolina Scarton

arXiv 2608.09510首次发表:更新:

发表机构

University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出BiBiR迭代框架测试虚假信息检测器鲁棒性,结合回译与LLM角色改写的攻破技术实现95%标签翻转率,DASS架构三元组对比模型准确率72.68%,优于基线模型。

AI 中文摘要

随着大语言模型(LLM)让大规模生成和改写误导性内容变得愈发容易,检测社交媒体上的机器生成虚假信息正变得越来越困难。静态基准评估仅在固定的保留数据集上测量检测器性能,无法捕捉到当帖子被刻意转换以逃避分类时检测器的表现。本文将“构建它、攻破它、修复它”框架调整为“构建它、攻破它、重复它(BiBiR)”:这是一种迭代会话,旨在在迭代对抗条件下对检测器的鲁棒性进行压力测试,评估当虚假信息帖子被系统地转换以逃避分类时,模型是否仍然可靠。在五次迭代中,研究发现,最佳对抗攻破者的转换方法结合了回译和基于LLM角色的改写,而表现最佳的技术达到了95%的标签翻转率(LFR),同时仍保留了原始帖子的含义。最佳构建者的模型是具有动态锚点切换(DASS)架构的三元组对比模型,该模型在最具鲁棒性的攻破者的对抗攻击集上实现了72.68%的平均准确率,比强基线(微调后的e5-small-LoRA)高出15个百分点。结果表明,迭代框架最能暴露检测器的弱点并推动鲁棒性改进;然而,它可能仍然需要语义保留分析来区分有效的对抗逃避与改变了原始虚假信息主张含义的转换。

英文摘要

Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on fixed held-out datasets, do not capture how detectors behave when posts are deliberately transformed to evade classification. This paper adapts the Build it, Break it, Fix it framework into Build it, Break it, Repeat (BiBiR): iterative sessions designed to stress-test detectors' robustness under iterative adversarial conditions, evaluating whether models remain reliable when disinformation posts are systematically transformed to evade classification. Across five iterations, the findings show that the best adversarial breakers' transformations came from a combination of back-translation and LLM persona-based rewriting, with the best performing technique achieving a 95% label flip rate (LFR), whilst still preserving the meaning of the original posts. The best builders' model was a triplet contrastive model with a dynamic anchor switching (DASS) architecture, which achieved an average accuracy of 72.68%, outperforming the strong baseline (a fine-tuned e5-small-LoRA) by 15 percentage points on the most robust set of breakers' adversarial attacks. The results demonstrate that an iterative framework best exposes detector weaknesses and pushes robustness improvements; however, it may still require semantic preservation analysis to distinguish valid adversarial evasion from transformations that changed the original disinformation claims' meaning.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑