arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

InfoOps Bench:实时信息行动安全基准

InfoOps Bench: A live information operations safety benchmark

Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright, John Gallacher

arXiv 2607.28503首次发表:更新:

AI 中文总结

该研究提出实时信息行动安全基准InfoOps Bench,测试17个前沿语言模型抵御国家支持信息行动的能力,发现多数模型可被利用,中国开发模型对涉华事实主张合规性显著下降,凸显模型可用性与安全性的平衡挑战。

AI 中文摘要

本文提出了一种主动、持续更新的AI基准,用于衡量前沿语言模型抵御被国家支持的信息行动利用的完整性。我们利用实时监测管道中的2100多个信息行动,该管道追踪俄罗斯、中国和伊朗国家支持的信息资产。随本文一同发布的还有一个配套网站,该网站每周更新国家支持媒体传播的最突出主张,网址为:this http URL。该基准的动态特性使其能够抵御饱和。在基准测试中,我们测试了来自8家提供商的17个模型,涵盖四种提示框架。我们发现大多数模型都可被用于信息行动。完整性得分定义为拒绝请求的百分比,范围为8.8%至94.5%,差距达85.7个百分点,且该差距与模型规模无关。模型选择还会改变产生的行动特征:一些模型编造细节,产生比源材料更有害的输出;另一些模型在符合要求的同时化解主张;事实核查率从2.9%到72.9%不等。抵御信息行动的完整性至少部分与即使面对良性主张也拒绝生成内容相关,这凸显了平衡模型可用性与安全性的挑战。除一个例外(this http URL的GLM 5.2),中国开发的模型在面对基于事实但针对中国的主张时,合规性大幅下降,与匹配的良性主张相比下降了48-70个百分点。

英文摘要

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state "information operations": intentional, coordinated activities by one state to influence public opinion and information ecosystems in another state. These information operations are a well documented, persistent threat against contemporary democracy. Our benchmark is based on real examples from over 2,100 information operations drawn from a live monitoring pipeline which tracks online information assets with links to authoritarian regimes. Alongside this paper, we also release a companion website that updates the benchmark weekly with new claims. The dynamic nature of this public facing benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations at least some of the time. Integrity scores, defined as the share of judged responses in which the model neither preserved nor amplified the claim, range from 9.3% to 91%, an 81.7-percentage-point spread not explained by model size. Models approach participation in information operations in a variety of ways. Some models fabricate details and produce output more harmful than the original input claim; others make claims less harmful even while complying and producing some output. Fact-checking rates vary from 3.2% to 80.8%. Integrity against information operations is at least partly related to refusal to produce content even for benign claims, illustrating the challenge of balancing model usability with safety. Overall, our results show the potential for contemporary information operations to be substantially aided by frontier AI models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑