VEX-Bench:评估LLM生成虚假信息的验证复杂度基准
VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
- The University of Melbourne(墨尔本大学)
- Deakin University(迪肯大学)
- Singapore Management University(新加坡管理大学)
- City University of Hong Kong(香港城市大学)
- Queensland University of Technology(昆士兰理工大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VEX-Bench评估LLM生成虚假信息的验证复杂度,提出VEX分数综合度量,发现LLM生成高复杂度虚假信息成本远低于验证,导致资源错配风险。
AI中文摘要:
大型语言模型(LLM)使得虚假信息的生成成本低廉,但验证成本并未降低,从而在信息生态系统中造成了日益加剧的不对称性。在时间、人力和预算紧张的情况下,媒体组织、平台和事实核查人员依赖筛选来确定优先验证的内容。我们引入了VEX-Bench,一个统一的基准,用于评估LLM生成的虚假信息在筛选过程中所感知的验证复杂度,涵盖多种模型和生成方法。验证复杂度沿多个维度进行评估,这些维度源自新闻实践和事实核查实践,包括可核查性、危害潜力、来源可信度信号、冒充合法性和预期验证工作量。我们将VEX分数定义为一个综合度量,结合了诱导产出和验证复杂度,以量化生成内容如何消耗有限的验证能力。我们构建了一个涵盖两类虚假信息、6个高风险领域和60个真实世界主题的基准,并评估了7个前沿LLM和7种生成方法,产生了5,880篇文章。我们采用LLM作为评判者进行可扩展评估,并使用内容分析方法进行验证,包括用于注释者间可靠性的序数Krippendorff α,辅以事实核查代理进行验证。我们的研究结果表明,没有单一方法在所有维度上占优,凸显了多维评估的必要性。LLM生成高VEX虚假信息的成本比基于代理的验证低3倍至169倍。此类内容通常在筛选中被优先考虑,消耗稀缺的验证资源,并在资源受限的验证系统中引入系统性资源错配风险。代码已公开在我们的GitHub仓库中。
英文摘要:
Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introduce VEX-Bench, a unified benchmark for evaluating the verification complexity of LLM-generated misinformation, as perceived during screening, across models and generation methods. Verification complexity is assessed along multiple dimensions derived from journalistic and fact-checking practices, capturing checkability, harm potential, source credibility signals, imposter legitimacy, and expected verification effort. We define the VEX score as an integrated measure combining elicitation yield and verification complexity to quantify how generated content consumes limited verification capacity. We construct a benchmark spanning two misinformation categories, 6 high-stakes domains, and 60 real-world topics, and evaluate 7 frontier LLMs and 7 generation methods, yielding 5{,}880 articles. We employ an LLM-as-judge for scalable evaluation and validate it using content-analysis methodology, including ordinal Krippendorff $α$ for inter-annotator reliability, complemented by fact-checking agents for verification. Our findings show that no single method dominates all dimensions, underscoring the need for multi-dimensional evaluation. LLMs can generate high-VEX misinformation at 3$\times$ to 169$\times$ lower cost than agent-based verification. Such content is often prioritized during screening, consuming scarce verification resources and introducing a systematic risk of misallocation in resource-constrained verification systems. The code is publicly available in our \href{https://github.com/HanxunH/VEX-Bench}{GitHub repository}.