FFSlim:一种用于多模态数据存储与检索的高效轻量格式
FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval
AI总结:
针对多模态数据存储检索的I/O瓶颈问题,提出轻量格式FFSlim,通过三个组件提升效率,实现高吞吐量并降低端到端训练时间
AI中文摘要:
随着大规模媒体-文本语料库的快速扩张,多模态数据集愈发需要高效的存储与检索能力。现有格式如Files、TDP和FFRecord在单模态数据场景下表现尚可,但在多模态环境中暴露出根本性局限,包括存储冗余、大量小文件开销、不利于缓存的布局以及繁重的索引结构。这些问题共同推高了存储与内存占用,使I/O成为实际训练工作负载中的主要瓶颈。本文提出FFSlim,一种用于存储和检索多模态数据的轻量格式。FFSlim通过三个组件提升存储效率与加载吞吐量:一是统一文件格式,可消除媒体重复并避免小文件泛滥;二是自适应检索机制,支持低开销的对级访问并加速重复媒体加载;三是冗余检测与聚合模块,可将现有数据集转换为FFSlim布局。实验结果表明,FFSlim的数据加载和写入吞吐量平均比最强基准高2.07倍和8.26倍,且存储与索引开销极小。因此,这些底层I/O加速使FFSlim在七种不同的多模态模型上,将端到端训练时间降低了5.36%-14.18%。
英文摘要:
With the rapid expansion of large-scale media-text corpora, multi-modal datasets increasingly require efficient storage and retrieval. Existing formats such as Files, TDP, and FFRecord work adequately for uni-modal data but expose fundamental limitations in multi-modal settings, including storage redundancy, massive small-file overheads, cache-unfriendly layouts, and heavy index structures. These issues jointly inflate storage and memory usage and make I/O the dominant bottleneck in real training workloads. We present FFSlim, a lightweight format for storing and retrieving multi-modal data. FFSlim improves storage efficiency and loading throughput through three components: a unified file format that removes media duplication and avoids small-file proliferation; an adaptive retrieval mechanism that enables low-overhead pair-level access and accelerates repeated media loading; and a redundancy detection and aggregation module that converts existing datasets into the FFSlim layout. The experimental results demonstrate that FFSlim achieves 2.07x and 8.26x higher data loading and write throughput on average than the strongest baseline, with minimal storage and index overhead. Consequently, these underlying I/O accelerations enable FFSlim to reduce end-to-end training time by 5.36%-14.18% across seven diverse multi-modal models.