RosettaBitcoin:关于智能体辅助共识验证器的验证基础设施的人工制品支持的经验报告
RosettaBitcoin: An Artifact-Backed Experience Report on Verification Infrastructure for Agent-Assisted Consensus Validators
浏览论文内容
中文总结 AI 辅助
本经验报告以 RosettaBitcoin 项目的 12 个 Bitcoin testnet4 共识验证器为研究对象,分析其验证基础设施的人工制品,提出明确的失败记录等可提升智能体辅助系统可审计性,需进一步实验验证该基础设施的作用。
中文摘要 AI 辅助
智能体辅助软件项目通常通过演示或综合基准进行报告,这些演示或基准掩盖了正确性主张是如何被认可的。本经验报告研究了 RosettaBitcoin,这是一个由单人开发者构建的项目,基于其不可变的 2026 年 6 月 17 日的软件快照(DOI:https://doi.org/10.5281/zenodo.20738249)构建了 12 个独立实现的 Bitcoin testnet4 共识验证器。我们分析了该快照的跟踪 SQLite 证据数据库、精心整理的人工制品索引、一致性夹具、验证脚本、阻塞记录和版本历史。在该快照中,所有 12 个移植版本都拥有移植专属的 45/45 脚本语料库证明和严格的 5000 个区块基线;其中 9 个拥有规范的干净 50000 个区块、100000 个区块以及 100000 之后的验证通道;Java 版本有一个耗时 19.86 秒的近 tip 维护人工制品。没有任何移植版本拥有从空状态到 tip 的证明,也没有任何移植版本满足项目的全二进制节点门槛;Docker 和活节点的能力差距仍然存在。人工制品历史还记录了一个耗时 3 小时 17 分 57 秒的 Zig 脚手架到 50000 跨度的记录,但观察到的间隔描述的是非等效任务,无法估算工作量、生产力或因果关系。一份独立的诊断补充材料保留了证据:纯 Mojo 加密后端验证了高度为 100000 的新状态并恢复到 140234,在 45 案例影子比较中达成一致,拒绝了 6 个精心制作的无效类,并被 3 个定向突变终止。该证据是非规范的、不可比较的且受类限制。该案例表明,明确的失败记录、夹具、移植专属证明和验证导入可以使智能体辅助系统更具可审计性,需要进行受控消融实验和外部复现,以测试此类基础设施是否会因果性地改善开发结果。
英文摘要
Agent-assisted software projects are often reported through demonstrations or aggregate benchmarks that conceal how correctness claims were admitted. This experience report studies RosettaBitcoin, a single-developer project that built twelve separately implemented Bitcoin testnet4 consensus validators, through its immutable 17 June 2026 software snapshot (DOI 10.5281/zenodo.20738249). We analyze the snapshot's tracked SQLite evidence database, curated artifact index, conformance fixtures, validation scripts, blocker records, and version history. At the snapshot, all twelve ports had port-owned 45/45 script-corpus proofs and strict 5,000-block baselines. Nine had canonical clean 50,000-block, 100,000-block, and post-100,000 validation lanes. Java had one 19.86-second near-tip maintenance artifact. No port had an empty-state-to-tip proof, and no port satisfied the project's binary full-node gate; Docker and live-node capability gaps remained. The artifact history also records a 3 h 17 min 57 s Zig scaffold-to-50,000 span, but the observed intervals describe non-equivalent tasks and cannot estimate effort, productivity, or causality. A separate diagnostic supplement preserves evidence that a pure-Mojo cryptographic backend validated fresh state to height 100,000 and resumed to 140,234, agreed on a 45-case shadow comparison, rejected six crafted invalid classes, and was killed by three targeted mutations. That evidence is noncanonical, noncomparable, and class-bounded. The case suggests that explicit failure records, fixtures, port-owned proofs, and validating imports can make agent-assisted systems more auditable. Controlled ablations and external replications are needed to test whether such infrastructure causally improves development outcomes.