发表机构
WMG, University of Warwick(华威大学制造工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型对夺旗赛的影响,通过混合方法绘制人机能力边界,发现社区对AI参赛分歧源于竞赛目的不明。为此提出含分层分组、抗LLM挑战设计等的保障框架及决策工具,其适用于网络安全中以结果证能力的场景。
AI 中文摘要
夺旗赛(CTF)是网络安全领域最有效的训练场之一,能培养密码学、网络攻击和二进制攻击等方面的实践技能。大语言模型(LLMs)如今能以最少人力投入解决越来越多的挑战,引发了关于公平性、排名有效性以及参与是否仍能带来有价值学习的紧迫问题。本文报告了一项关于大语言模型对现代CTF影响的混合方法研究,综合已发布的基准测试(包括近期政府评估)、三类现场竞赛的案例研究、对社区讨论人工智能使用的公共渠道的结构化观察,以及对经验丰富的参与者和组织者的半结构化访谈。我们按类别绘制了当前人机能力边界,发现社区对于是否允许使用人工智能的分歧,源于一个未明确的先验问题:竞赛的目的是什么。在此背景下,我们提出了一个由四部分组成的保障框架,包括分层竞赛分组、抗大语言模型的挑战设计、用于调查的遥测技术以及社区行为准则草案,还有一个将保障措施组合与竞赛既定目的相联系的决策工具。该论点不仅适用于CTF,还适用于网络安全中任何以展示结果作为潜在能力证据的场景。
英文摘要
Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, and binary exploitation are now reliably automated while narrower sub-categories continue to resist. We find that community disagreement about whether AI should be permitted is downstream of an undeclared prior question: what a competition is for. Against this backdrop we contribute a four-component safeguard framework, combining tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively, and a draft community code of conduct, together with a decision tool that ties the combination of safeguards to a competition's declared purpose. The argument reaches beyond CTFs to any setting in cybersecurity where a demonstrated result is taken as evidence of an underlying ability.
Comments20 pages, 3 figures, 1 table