发表机构
Institute of Information Engineering, Chinese Academy of Sciences; University of Illinois Chicago; State Key Laboratory of Media Convergence and Communication, Communication University of China; University of Chinese Academy of Sciences(中国科学院信息工程研究所; 伊利诺伊大学香槟分校; 中国传媒大学媒体融合与传播国家重点实验室; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型在错误信息生态系统中的问题,引入角色层框架统一风险与防御,组织相关攻击、研究检测与验证方法、分析漏洞及讨论对策,指出从静态检测到风险评估、强化验证管道、部署可审计人工参与验证系统这三个关键挑战。
AI 中文摘要
大语言模型(LLMs)已将错误信息从主要以内容为中心的问题转变为更广泛的生态系统级安全挑战。被滥用时,LLMs产生的风险超出虚假内容生成,能攻击错误信息防御所依赖的社会背景、证据来源、检索语料库和验证工作流程。本文引入角色层框架统一这些风险与防御。角色维度将LLMs表征为攻击者、防御者和验证系统的脆弱组件,层维度涵盖内容、社会背景、证据环境和验证工作流程。基于此框架,组织基于LLMs的攻击,研究基于LLMs的检测与验证方法,分析以LLM为中心的检测范式中的漏洞,讨论现有对策。在此基础上,识别出三个关键开放挑战:从静态检测精度转向预算生态系统级风险评估,强化以LLM为中心的验证管道抵御对抗性操纵,部署可审计的人工参与验证系统用于可靠的现实世界错误信息防御。
英文摘要
Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level security challenge. When misused, LLMs create risks beyond false content generation, enabling attacks on the social contexts, evidence sources, retrieval corpora, and verification workflows that misinformation defense depends on. In this paper, we introduce a role-layer framework to unify these risks and defenses. The role dimension characterizes LLMs as attackers, defenders, and vulnerable components of verification systems, while the layer dimension covers content, social contexts, evidence environments, and verification workflows. Guided by this framework, we organize LLM-enabled attacks, investigate LLM-based detection and verification methods, analyze vulnerabilities in LLM-centric detection paradigms, and discuss existing countermeasures against LLM-enabled attacks. Building on this synthesis, we identify three key open challenges: moving from static detection accuracy to budgeted ecosystem-level risk evaluation, hardening LLM-centered verification pipelines against adversarial manipulation, and deploying auditable human-in-the-loop verification systems for trustworthy real-world misinformation defense.
Comments35 pages, 8 figures