发表机构
Georgia Tech; Dolby Labs(佐治亚理工学院; 杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LayoutBench针对多媒体数据云存储布局开展基准测试,评估三种布局策略,揭示其在不同检索场景下的性能差异,为云存储布局选择提供数据支撑。
AI 中文摘要
现代多媒体机器学习工作负载日益将大规模数据集存储在AWS S3等云对象存储服务中,样本在存储中的物理组织方式(即存储布局)直接影响其检索速度与成本。然而,当前用于指导存储决策的基准测试聚焦于数据库引擎和查询处理,均未系统评估不同存储布局对多媒体数据检索的性能。我们提出LayoutBench,这是首个填补该空白的基准测试,它评估三种代表性布局策略:将每个样本存储为单个对象(L1)、按顺序将样本打包为tar归档文件(L2)、将样本组织为Parquet文件中的列(L3)。我们在覆盖不同网络带宽和内存层级的6种AWS EC2实例配置上,针对ImageNet数据集的11种不同结果集大小的查询,测量检索时间、传输数据量和资金成本。实验显示,L2通过连接复用实现比L1和L3更低的延迟,但在检索规模极大时会失去该优势;L3在超大规模检索中速度最快,但因行组粒度在所有查询规模下传输的数据量显著更多,且需要多得多的内存;在所有布局中,数据传输成本主导总支出,L3的成本比L1或L2高一个数量级。
英文摘要
Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3. How these samples are physically organized in storage (i.e.,storage layout) directly affects how quickly and cheaply they can be retrieved. Yet the benchmarks used to guide storage decisions today focus on database engines and query processing, and none systematically evaluates how different storage layouts perform for multimedia data retrieval. We present LayoutBench, the first benchmark designed to fill this gap. It evaluates three representative layout strategies: storing each sample as an individual object (L1), sequentially packing samples into tar archives (L2), and organizing samples as columns in Parquet files (L3). We measure retrieval time, data transferred, and monetary cost using 11 queries of varying result-set sizes on ImageNet across six AWS EC2 instance configurations that span different network bandwidth and memory tiers. Our experiments reveal that L2 achieves lower latency than L1 and L3 through connection reuse, but loses this advantage as retrieval sizes become very large. L3 is the fastest for very large retrievals but transfers substantially more data across all query sizes due to row-group granularity, and requires significantly more memory. Across all layouts, data transfer cost dominates total expenditure, with L3 costing an order of magnitude more than L1 or L2.