arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

tse_tick:用于解析和查询东京证券交易所日经NEEDS逐笔成交数据的Python库

tse_tick: A Python Library for Parsing and Querying Nikkei NEEDS Tick Data from the Tokyo Stock Exchange

Kazumi Li, Masataka Hayashi, Teruo Nakatsuma, Peter Romero

arXiv 2608.23053首次发表:更新:

AI 中文总结

tse_tick是一款开源Python库,可将东京证券交易所日经NEEDS的逐笔成交数据转换为可查询的结构化格式,其工程优化实现了比pandas更快的解析与查询速度,已发布于PyPI。

AI 中文摘要

东京证券交易所的逐笔成交与报价数据通过日经NEEDS服务分发,形式为数千个ZIP压缩CSV归档文件,涵盖四种数据类型,具有随时代变化的模式和日语布局。我们提出tse_tick,这是一个开源Python库,可将这些原始归档文件转换为干净的类型化Polars DataFrame和可通过DuckDB查询的Hive分区Parquet存储。该库提供两条访问路径,共享一个解析与清洗核心:一条是一次性读取器,可直接从原始ZIP文件返回按股票代码和时间过滤的DataFrame;另一条是两阶段的导入-查询管道,具备可恢复、内存感知的并行导入、部件剪枝以及用于行组剪枝的物化日内时间键。该库的贡献更多在于工程而非解析:导入按每日原子单元运行,其完成情况由覆盖标记记录,而非从文件存在性推断;写入以有界块流方式进行,因此峰值内存与交易日规模无关(实测最差交易日从24.5 GB降至2.4 GB);内存感知进程池会根据可用内存调整自身大小;部件剪枝仅打开股票代码可占据的归档部件。所有四种数据类型均附带完整的英语和日语列定义,且存在转换层将yfinance、Polygon和ccxt的名称映射到tse_tick的等效名称。在普通16线程工作站上的基准测试中,解析代表交易日9个部件之一的480万行归档部件,速度比原始pandas原型快59.8倍(与引擎匹配的pandas基线相比快34.3倍),从存储中查询单个股票代码的时间窗口比扫描等效CSV的pandas快约410倍。tse_tick可在PyPI上获取(通过pip install tse-tick安装),遵循MIT许可证。

英文摘要

Tick-level trade-and-quote data for the Tokyo Stock Exchange is distributed through the Nikkei NEEDS service as thousands of zipped CSV archives spanning four data types with era-dependent schemas and Japanese-language layouts. We present tse_tick, an open-source Python library that converts these raw archives into clean, typed Polars DataFrames and a Hive-partitioned Parquet store queryable through DuckDB. The library offers two access paths sharing one parse-and-clean core: a one-shot reader that returns a ticker- and time-filtered DataFrame directly from raw ZIP files, and a two-stage ingest-then-query pipeline with resume-safe, memory-aware parallel ingestion, part-pruning, and a materialized intraday time key for row-group pruning. The engineering, more than the parsing, is what the library contributes: ingestion runs in per-date atomic units whose completion is recorded by coverage markers rather than inferred from file existence, writes stream in bounded morsels so that peak memory is independent of trading-day size (24.5 GB to 2.4 GB on the worst measured day), a RAM-aware process pool sizes itself to available memory, and part-pruning opens only the archive parts a ticker can occupy. Full English and Japanese column definitions ship for all four types, and a translation layer maps yfinance, Polygon, and ccxt names onto their tse_tick equivalents. In benchmarks on a commodity 16-thread workstation, parsing a representative 4.8-million-row archive part, one of a trading day's nine parts, is 59.8x faster than the original pandas prototype (34.3x against an engine-matched pandas baseline), and a single-ticker time-window query from the store completes roughly 410x faster than a pandas scan of the equivalent CSV. tse_tick is available on PyPI (pip install tse-tick) under the MIT license.

Comments23 pages, 2 figures. Code: https://github.com/tse-tick/tse_tick

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑