打造我自己的验证器调用清洁的证据包的四种方法:证据包验证器的拒绝点变异测试
Four Ways to Forge a Bundle My Own Verifier Calls Clean: Refusal-Site Mutation Testing of an Evidence-Bundle Verifier
浏览论文内容
中文总结 AI 辅助
本文对证据包验证器开展拒绝点变异测试,发现其存在空通过缺陷,经测量、修复后提升了验证器的可靠性,所有数据均来自作者自建系统,唯一外部数据为指定审计结果。
中文摘要 AI 辅助
我构建了一个协议,其前提是陌生人可以离线重新运行我的声明并得到相同的答案。一名外部工程师对其进行审计并破解了它:一个标题数字为假的证据包被验证为清洁,最廉价的伪造品仅需四个字节。我合并了他的修复方案,随后用我自己的工具对修复后的验证器进行检测,在他的审计未覆盖的地方又发现了四处相同的缺陷,其中最廉价的仅需一个大写字母。统一的缺陷并非密码学或奇异问题:一条检查在未检查任何内容的路径上报告成功。“空通过”是一个可行的标签,而非一项发现;第4节指出已有文献占据了该领域。因此我停止收集轶事并进行测量:在f59fb62版本中,根据第6节的提取规则,该验证器暴露了112个拒绝点;删除其中75个后,整个套件及所有篡改夹具仍保持为绿色,得分为0.330。手动发现的四个伪造品中的三个属于存活类别;第四个是无拒绝点的义务。单独评分时,为证明验证器可以弃权(不执行)而构建的16个夹具语料库捕获了10个。对拒绝点本身进行测试后,得分升至0.941,随后在92e4548版本中,基于146个站点的增长种群得分达到1.000;这些分母不同且中间序列非单调,因此第7.5节包含全部11个,而非仅表2的5行。修复四个发现的缺陷后,112个站点变为119个中的39个,原有站点仍为37个。研究期间,我的测量工具七次在未测量任何内容时报告成功;其中四个是为检测此类情况而构建的,一个返回了完美的1.000。此处的所有数字均是我在自己编写的系统上自我测量得到的,基于我自己仓库的闭环注册表;唯一的外部数据点是第2.2节的审计结果。此处明确说明而非隐藏:这是论文的可信度,而非警告。
英文摘要
I built a protocol whose premise is that a stranger can re-run my claims offline and get the same answer. An outside engineer audited it and broke it: a bundle whose headline numbers were false verified clean, the cheapest forgery four bytes. I merged his fix, then pointed my own instruments at the fixed verifier and found the same defect four more times, in places his audit did not reach. The cheapest is one capital letter. The unifying defect is not cryptographic or exotic: a check that reports success along a path where it never examined anything. Vacuous pass is a working label, not a discovery; Section 4 names the literatures already occupying it. So I stopped collecting anecdotes and measured. At f59fb62, under the extraction rule of Section 6, the verifier exposes 112 refusal sites; 75 could be deleted with the whole suite and every tamper fixture still green, a score of 0.330. Three of the four hand-found forgeries fall in surviving classes; the fourth is an obligation with no refusal site. Scored alone, the sixteen-fixture corpus built to prove the verifier can refuse catches 10. Testing the refusals themselves took it to 0.941, then to 1.000 at 92e4548 over a grown population of 146 sites; those denominators differ and the series between them is non-monotone, so Section 7.5 carries all eleven, not just the five rows of Table 2. Fixing the four found defects instead moved 37/112 to 39/119, leaving the pre-existing sites at 37. Seven times during this study my own measuring tools reported success while measuring nothing; four were built to detect this class, and one returned a perfect 1.000. Every number here is self-measured on a system I wrote, over a registry that is a closed loop of my own repositories; the one external data point is the audit of Section 2.2. That is stated here rather than buried: it is the paper's credibility, not a caveat.