发表机构
Theta Labs; Starknet Foundation; Brevis; MultiVM Labs; StarkWare; Adam Mickiewicz University Poznań; Warsaw University of Technology; Octav; Pauli Group; ScienceVR; Sei Labs; Stanford Free Systems Lab; Eigen Labs; Ethereum Foundation(Theta Labs; Starknet基金会; Brevis; MultiVM实验室; StarkWare; 波兹南亚当·密茨凯维奇大学; 华沙理工大学; Octav; Pauli集团; ScienceVR; Sei实验室; 斯坦福自由系统实验室; Eigen实验室; 以太坊基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出开放自动研究范式,通过公共排行榜优化Shor算法中secp256k1点加法电路,将时空评分降低86.1%,最佳电路使用1,151量子比特,评分比谷歌阈值低50%以上。
AI 中文摘要
我们提出开放自动研究(Open Autoresearch)范式,在该范式中,人类与AI智能体将经评估者验证的改进发布到公共排行榜上。我们在http URL中实例化该范式,优化可逆的secp256k1点加法电路,这是椭圆曲线密码学中Shor算法的一个瓶颈。该基准最小化受时空启发的评分$S=Q\times T$,其中$Q$为峰值逻辑量子比特宽度,$T$为平均执行的Toffoli门数量。参与者将$S$降低了86.1%。在数据截止日期(2026年7月26日),得分最高的电路使用1,151个量子比特和1,299,453个平均执行的Toffoli门,得到$Q\times T\approx14.96$亿。这比谷歌公布的点加法评分阈值(arXiv:2603.28846)低50%以上,尽管采用了不同的核算约定。由于基准在经典侧提供一个加数,我们构建了一个兼容窗口加法的相干变体,实现窗口化Shor所需的单次调用接口。该变体使用1,162个量子比特和1,684,161个平均执行的Toffoli门。在100,000个随机输入上,其实验成功概率为$\hat{p}=0.99809$,在独立可重跑的每次调用敏感性模型下,$Q\times T/\hat{p}\approx19.61$亿,但这并非完整的Shor成功估计。其量子比特和Toffoli计数均低于谷歌公布的阈值和Schrottenloher报告的工作点(arXiv:2606.02235),尽管不同的接口、核算约定和验证范围排除了形式上的支配关系。在截止日期后,评分进一步降至12.59亿,同时一个独立的低宽度电路达到813个量子比特。公开记录显示AI智能体补充了人类判断,为在可高效评估、机器可验证的目标上进行开放自动研究提供了证据。
英文摘要
We propose Open Autoresearch, a paradigm in which humans and AI agents publish evaluator-verified improvements to a public leaderboard. We instantiate it in ECDSA.Fail, optimizing reversible secp256k1 point-addition circuits, a bottleneck in Shor's algorithm for elliptic-curve cryptography. The benchmark minimizes the spacetime-inspired score $S=Q\times T$, where $Q$ is peak logical qubit width and $T$ is average executed Toffoli count. Participants reduced $S$ by 86.1%. At the data cutoff (26 July 2026), the best-scoring circuit uses 1,151 qubits and 1,299,453 average executed Toffoli gates, giving $Q\times T\approx1.496$ billion. This is more than 50% below Google's published point-addition score thresholds (arXiv:2603.28846), under different accounting conventions. Because the benchmark supplies one addend classically, we construct a coherent windowed-addition-compatible variant implementing the single-call interface required by windowed Shor. It uses 1,162 qubits and 1,684,161 average executed Toffoli gates. On 100,000 random inputs, its empirical success probability is $\hat{p}=0.99809$, giving $Q\times T/\hat{p}\approx1.961$ billion under an independently rerunnable per-call sensitivity model, not a full-Shor success estimate. Its qubit and Toffoli counts lie below Google's published thresholds and Schrottenloher's reported operating points (arXiv:2606.02235), although differing interfaces, accounting conventions, and validation scope preclude formal dominance. After the cutoff, the score was further reduced to 1.259 billion, while a separate low-width circuit reached 813 qubits. The public record shows AI agents complementing human judgment, providing evidence for open autoresearch on efficiently evaluable, machine-checkable objectives.
Comments62 pages, 10 figures. Project website and latest results: https://ecdsa.fail and source code: https://github.com/Layr-Labs/ecdsafail-challenge