AI 中文总结
本研究开发了开源模块化基准测试平台,集成评估DNA存储编解码器,量化其多维度权衡,为DNA数据存储的编解码器选择与应用提供标准化评估框架。
AI 中文摘要
背景:脱氧核糖核酸(DNA)数据存储是一种具有超密集和持久信息保存潜力的范式。然而,编码方案(或称编解码器)的快速增殖,每种都有各自的设计约束和报告实践,导致了碎片化的格局,缺乏标准化的比较评估。方法:我们开发了一个开源、模块化的基准测试平台,该平台系统地集成并评估了最先进的DNA存储编码和解码方法(编解码器)。我们的方法使用一组经过精心筛选的多样化基线数据,并应用与DNA数据存储联盟共识标准一致的多维度评估标准,这些标准包括编码/解码吞吐量、计算效率、针对替换、插入和删除的纠错性能,以及成本效率。结果:所开发的平台集成了标准化的编码和解码包装函数,允许集成新方法,并通过全面的可视化和表格报告实现可重复的自动化评估。使用默认参数和多个指标对当代和经典编解码器进行基准测试表明,没有任何一种算法在所有评估维度上都是最优的。信息密度、成功率、运行时间和成本之间的权衡被量化,并显示为面向未来的格式设计中的关键因素。结论:我们的工作建立了一个严格标准化的开源评估框架,该框架支持可重复的基准测试,支持基于证据的编解码器选择,并为将DNA数据存储从实验研究转化为可部署的归档系统提供了必要基础。
英文摘要
Background: Deoxyribonucleic acid (DNA) data storage is a paradigm with great potential for ultra-dense and durable information preservation. However, the rapid proliferation of coding schemes, or codecs, each with their own design constraints and reporting practices, has led to a fragmented landscape that lacks a standardized comparative assessment. Methods: We developed an open-source, modular benchmarking platform that systematically integrates and evaluates state-of-the-art DNA storage encoding and decoding methods (codecs). Our approach uses a curated, diverse set of baseline data and applies multidimensional assessment criteria that are aligned with the consensus standard of the DNA Data Storage Alliance. These criteria include encoding/decoding throughput, computational efficiency, error correction performance across substitutions, insertions, and deletions, and cost efficiency. Results: The developed platform integrates standardized wrapper functions for encoding and decoding, allows for the integration of new methods, and automates reproducible evaluations with comprehensive visual and tabular reporting. Benchmarking both contemporary and classical codecs using their default parameters and multiple metrics demonstrates that no single algorithm is optimal across all evaluated dimensions. The trade-offs between information density, success rate, runtime, and cost are quantified and shown to be critical factors in the design of future-proof formats. Conclusions: Our work establishes a rigorously standardized, open-source evaluation framework that enables reproducible benchmarking, supports evidence-based codec selection, and provides the necessary foundation for translating DNA data storage from experimental research into deployable archival systems.
Comments38 pages, 12 figures