智能压缩:从湖表元数据预测压缩效用
Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata
浏览论文内容
中文总结 AI 辅助
该研究针对开放湖表格式的小文件问题,提出模拟框架与XGBoost模型预测压缩效用,发现单一分区阈值可做二元压缩决策,验证了泛化性并揭示压缩对不同查询的影响,代码数据公开。
中文摘要 AI 辅助
开放湖表格式会随时间积累大量小数据文件,这会降低查询性能。判断何时值得进行压缩仍依赖阈值驱动的方式,但哪些元数据特征实际决定了压缩效用尚不明确。我们提出了一个开放模拟框架,生成了2376个Apache Iceberg表,文件大小跨度达三个数量级,在不读取数据的情况下从清单文件中提取17个元数据特征,并训练XGBoost模型来预测连续文件缩减率(R²=0.998,RMSE=0.013)。二元压缩决策可通过单一分区级阈值max_files_per_partition>4轻易区分,无需使用训练好的模型。对96个TPC-H表进行的跨模式验证证实了无需重新训练即可泛化(R²=0.976)。查询基准测试显示,压缩对元数据密集型查询有益,但会因降低任务并行度而减慢全扫描聚合的速度。所有代码和数据均公开可用。
英文摘要
Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains threshold-driven, but which metadata features actually determine compaction utility is not well understood. We present an open simulation framework that generates 2,376 Apache Iceberg tables spanning three orders of magnitude in file size, extracts 17 metadata features from manifest files without reading data, and trains XGBoost to predict the continuous file-reduction ratio (R2 = 0.998, RMSE= 0.013). The binary compaction decision turns out to be trivially separable by a single partition-level threshold max_files_per_partition> 4, requiring no learned model. Cross-schema validation on 96 TPC-H tables confirms generalisation without retraining (R2 = 0.976). A query benchmark reveals that compaction benefits metadata-heavy queries but can slow full-scan aggregations by reducing task parallelism. All code and data are publicly available.
发表机构
- European Central Bank(欧洲中央银行)
- European System of Central Banks(欧洲中央银行体系)
- IBM(国际商业机器公司)
机构由 AI 辅助整理,请以论文原文为准。