发表机构
University of Science and Technology of China; Huawei; South China University of Technology; Shanghai Jiao Tong University(中国科学技术大学; 华为; 华南理工大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态AI数据集处理速度不足的问题,推出专为AI数据集设计的VersaDB,通过页式存储、B+树索引等技术实现高效数据处理,可最高5.35倍加速且并行度适配性好。
AI 中文摘要
AI领域发展迅速,涌现出大量各类AI训练数据集,这些数据集包含文本、图像、音频等不同模态,且可能采用多种数据存储格式。随着AI硬件的进步,GPU、TPU、NPU等AI计算单元可大幅加快AI模型的训练速度,这反过来对更快的数据处理提出了更高需求。当使用现有AI处理框架处理不同模态和存储格式的数据集时,由于数据布局及用户处理数据的方式等问题,处理速度可能未达最优。因此,使用统一数据库存储多种数据格式可更好地管理和优化数据访问。本文介绍VersaDB,这是一款专为各类模态AI数据集设计的数据库。我们实现了基于页的存储系统,将结构化与非结构化数据分离;还生成了基于B+树的索引文件以加速数据访问。VersaDB支持自动分片,并维护分层元数据管理系统,在页、分片、全局层面均维护对应元数据,构成数据库高效运行的基础。我们还注重易用性,提供了将数据集直接转换为VersaDB的API,以及将CSV、TFRecord、.bin等流行AI数据存储格式转换为该数据库的API。实验表明,使用VersaDB可实现最高5.35倍的加速,并在不同并行度下保持一致的性能。
英文摘要
The AI field has been rapidly developing, leading to the emergence of a large number of AI training datasets of various types. These datasets contain different modalities, including text, images, audio, etc., and may come in various data storage formats. With the advancement of AI hardware, AI computation units like GPUs, TPUs, and NPUs can greatly accelerate the training speed of AI models, which in turn increases the demand for faster data processing. When using existing AI processing frameworks to handle datasets with different modalities and storage formats, processing speeds may be suboptimal due to issues such as data layout and the way users handle the data. Therefore, using a unified database to store multiple data formats can better manage and optimize data access. In this paper, we introduce VersaDB, a database designed specifically for AI datasets with various modalities. We implemented a page-based storage system, separating structured and unstructured data. Additionally, we generated B+ tree-based index files to accelerate data access. VersaDB supports automatic sharding and maintains a hierarchical metadata management system, with corresponding metadata maintained at the page, shard, and global levels, forming the foundation for the efficient operation of the database. We also focused on ease of use by providing APIs for directly converting datasets into VersaDB, as well as APIs for converting popular AI data storage formats (e.g., CSV, TFRecord, .bin) into VersaDB.Our experiments show that using VersaDB can achieve up to 5.35x acceleration and maintain consistent performance across different parallelism levels.