arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非对称金融数据集的生命周期感知归档:一项生产研究

Lifecycle-Aware Archival for Asymmetric Financial Datasets: A Production Study

Tulika Manek

arXiv 2608.12367首次发表:更新:

AI 中文总结

该研究针对金融交易数据库的存储效率与操作新鲜度矛盾,提出生命周期感知归档系统,解决名人分区问题,通过ID单调性去重技术实现热存储缩减、成本降低及性能提升。

AI 中文摘要

大规模金融交易数据库面临着操作新鲜度需求与存储效率之间的根本矛盾。我们介绍了Razorpay旗下一款金融交易服务的生命周期感知归档系统的设计、实现与生产评估,该系统在PostgreSQL Aurora(14版)上管理数十亿条记录,占用数十TB存储空间,并维持每秒数千次事务的峰值写入吞吐量。我们作出两项贡献:首先,我们分析性地定义了“名人分区问题”:生命周期状态分区将所有操作活跃行集中在单个默认分区中,导致每次状态转换时产生O(N×M)的规划开销、O(N)的执行I/O退化以及写入放大。其次,我们提出一种基于ID单调性的去重技术,该技术利用单调递增ID方案(Snowflake ID、ULID及等价方案)中的时间编码,仅将可能需归档的重新插入操作路由至温数据库查找,无需布隆过滤器、幂等性表或外部依赖。我们报告了已全面部署系统的生产结果:热存储容量减少95%、月度数据库基础设施成本降低53%、服务级p99处理延迟降低约60%、写入器CPU利用率降低51个百分点,且在未对主事务表进行架构变更的情况下维持了峰值写入TPS的运行。

英文摘要

Large-scale financial transaction databases face a fundamental tension between operational freshness requirements and storage efficiency. We present the design, implementation, and production evaluation of a lifecycle-aware archival system for a financial transaction service at Razorpay managing billions of records on PostgreSQL Aurora (version 14), occupying tens of terabytes of storage and sustaining peak write throughput in the thousands of TPS. We make two contributions. First, we analytically characterize the Celebrity Partition Problem: lifecycle-state partitioning concentrates all operationally active rows in a single default partition, incurring O(N x M) planning overhead, O(N) execution I/O regression, and write amplification on every state transition. Second, we present an ID-monotonicity-based deduplication technique that exploits temporal encoding in monotonically increasing ID schemes (Snowflake IDs, ULIDs, and equivalent) to route only potentially-archived re-inserts to a warm database lookup, requiring no Bloom filters, idempotency tables, or external dependencies. We report production results from a fully deployed system: 95% hot storage size reduction, 53% reduction in monthly database infrastructure cost, ~60% reduction in service-level p99 processing latency, 51 percentage-point reduction in writer CPU utilization, and sustained operation at peak write TPS - without schema changes to the primary transaction table.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑