arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Scratchy:面向EasyCrypt密码学证明生成的视觉草稿多模态推理

Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in EasyCrypt

Yupeng Ren, Zhaoxuan Li, Rui Zhang

arXiv 2609.06226首次发表:更新:

发表机构

Institute of Information Engineering, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院信息工程研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Scratchy,一种视觉草稿多模态方法,通过将证明对象规范化为类型化证明关系图并转换为视觉证明状态,引导多模态模型生成EasyCrypt密码学证明,实验表明该方法显著提升LLM性能。

AI 中文摘要

大语言模型(LLMs)近期在形式化证明生成方面取得了实质性进展,但在密码学领域仍面临独特挑战。计算安全性论证认为,有效证明必须协调概率、对抗游戏、不变量、假设和界,这些可由名为EasyCrypt的机器检查框架提供。尽管所有对象可能出现在可用上下文中,LLMs仍然难以处理,因为证明论依赖通常在线性表示中是隐式的,并分布在多个程序中。因此,本文提出Scratchy,一种视觉草稿方法,用于暴露这些依赖以进行多模态生成。给定自然语言安全性描述,连同形式化上下文和目标命题,证明对象可被规范化为类型化证明关系图。然后,一个保持结构的视觉编译器将该图转换为富含公式的视觉证明状态,引导多模态模型生成EasyCrypt证明。此外,引入了Scratchy-eval,一个源自可靠官方EasyCrypt文件的114任务数据集。它包含64个安全形式证明生成和50个多项选择知识测试。经过一系列评估,涵盖语义基础、关系不变量和游戏归约,经典LLMs如GPT-5.6-Sol和Claude-Opus-5从Scratchy的结构化视觉证明状态中获得了明显优势。这一对比表明,显式证明结构可以带来改进,多模态证明状态表示是计算机辅助密码学的一个有前景的方向。

英文摘要

Large language models (LLMs) have recently made substantial progress in formal proof generation, yet presenting distinctive challenges in cryptographic area. Computational security arguments posit that a valid proof must coordinate probability, adversarial games, invariants, assumptions and bounds, which can be provided by a machine-checked framework named EasyCrypt. Although all objects may appear in available context, LLMs still struggle because proof-theoretic dependencies are typically implicit in a linear representation and distributed across multiple programs. So, this paper presents Scratchy, a visual-scratchpad approach that exposes these dependencies for multimodal generation. Given the natural-language security description, with formal context and target propositions, the proof objects can be normalized into a typed proof-relation graph. Then a structure-preserving visual compiler transforms the graph into the formula-rich visual proof state that guides a multimodal model in generating the EasyCrypt proof. Also, the Scratchy-eval, a 114-task dataset derived from reliable official EasyCrypt files, has been introduced. It contains 64 security-form proof generations and 50 multiple-choice knowledge tests. After a series of evaluations, covering semantic grounding, relational invariants, and game reductions, classical LLMs like GPT-5.6-Sol and Claude-Opus-5 have gained a clear advantage from Scratchy's structured visual proof states. This contrast suggests that explicit proof structure can make the improvement and multimodal proof-state representation as a promising direction for computer-aided cryptography.

CommentsFirst version; 12 pages, 5 figures, and 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑