arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

密码分析基准测试:大语言模型能进行密码分析吗?

CryptanalysisBench: Can LLMs do Cryptanalysis?

Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski, Milad Nasr, Orr Dunkelman, Eyal Ronen, Florian Tramèr

arXiv 2607.18538首次发表:更新:

AI 中文总结

研究探讨大语言模型能否进行密码分析,引入包含六个密码原语家族191个任务的CryptanalysisBench基准测试,五个前沿模型在不同层级有一定破解率,不仅得出已知结果还产生新成果,发布该基准测试以追踪人工智能密码分析进展及压力测试候选方案。

AI 中文摘要

密码分析是寻找针对加密方案攻击的任务,处于数学推理和网络安全的交叉领域,而大语言模型在这两个领域发展迅速。密码分析既是前沿推理的纯净测试平台,又因研究的原语支撑数字安全而关乎重大。本文探讨大语言模型能否进行密码分析,答案日益肯定。我们引入密码分析基准测试(CryptanalysisBench),它包含六个密码原语家族的191个任务,主要来自四个美国国家标准与技术研究院(NIST)标准化竞赛。该基准测试有三层:已知有实际破解方法的原语;尚无已知实际破解方法的原语,在全强度和缩小版本下评估;密码分析前沿的生产原语挑战集。五个前沿模型在第一层方案中破解率达65%-86%,全强度下在第二层方案中破解6-12个,在所有缩小版本中破解24-61个。模型不仅得出已知结果,还产生了新的密码分析成果,如利用SpoC AEAD设计缺陷的密钥恢复攻击以及发现KINDI已发表的CCA安全证明中的错误。我们发布密码分析基准测试作为工具,以帮助追踪人工智能密码分析是否(或何时)成为重要因素,并作为在部署前对候选方案进行压力测试的框架。基准测试已揭示的攻击是快速发展前沿的早期快照,可能很快赶上并在某些方面超越已发表的现有技术水平。

英文摘要

Cryptanalysis - the task of finding attacks against cryptographic schemes - sits at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced fastest. Cryptanalysis represents both a clean testbed for frontier reasoning (as practical attacks can be automatically verified) and a domain with unusually high stakes, since the primitives under study underpin our digital security. In this paper we ask whether LLMs can do cryptanalysis, and find that the answer is increasingly yes. We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions. Our benchmark consists of three tiers: (i) primitives with known practical breaks; (ii) primitives with no known practical break, evaluated both at full strength and as scaled-down variants; and (iii) a challenge set of production primitives at the frontier of cryptanalysis. Five frontier models (Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and the open-weights GLM 5.2) break 65%-86% of Tier 1 schemes, 6-12 Tier-2 schemes at full strength, and 24-61 across all scaled-down variants. Beyond deriving known results, models produce novel cryptanalysis, such as a key-recovery attack that exploits a design flaw in the SpoC AEAD and an error in KINDI's published CCA-security proof, both to the best of our knowledge not previously known. We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment. The attacks that the benchmark already surfaces are an early snapshot of a fast-moving frontier that may soon match, and in places exceed, the published state of the art.

Comments46 pages, 5 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑